The change we reverted as a regression was just noise
Goal
You repeat the same target five times to measure the size of the noise first, measure a version with slightly increased latency and a version with greatly increased latency five times each, and judge by a rule which exceeds the noise. You turn that rule into a CI gate script that takes a baseline directory and a new measurement directory and emits an exit code, and confirm in numbers that a single-run comparison turns even the same code into a regression.
Why it matters
Judging a performance regression is not a statistics problem but an operations problem. If you set the threshold wrong, one of two things happens — false alarms are frequent so people turn the gate off, or the threshold is loose so real regressions slip through. Both cases become the same as having no gate. So the threshold must be not a percentage decided in a meeting room but noise actually measured on that machine at that time. To measure noise, you must run the same test several times without changing anything, and that naturally reveals that 'how many times to run' is a trade-off between time and sensitivity. If you run it once, the threshold becomes 0 and anything is a regression, and the more you run, the more confidently you can catch smaller regressions. Leaving this judgment not to human eyes but to a script's exit code is the last step of this lab.
Steps
- Start
/opt/lab/lt/lt-regression/svc.pyon 127.0.0.1:8080 withVERSION=base DELAY_MS=20 JITTER_MS=20 SEED=1234. Runhey -n 200 -c 5 -o csv http://127.0.0.1:8080/workfive times and save the outputs as/root/lt-regression/base/run1.csvthroughrun5.csv. For each run, compute p50 and write it as five lines to/root/lt-regression/01-noise.tsv— each line has two tab-separated columns,run<번호> <p50 밀리초>(the run number and p50 in milliseconds), with milliseconds to three decimal places. p50 is determined by converting the first column of the csv (response-time, in seconds) to milliseconds, sorting in ascending order, and taking theint(n/2)+1th value (counting from 1). Then write four lines to/root/lt-regression/01-spread.txt:min=,max=,spread_ms=<max - min>, andspread_pct=<(max - min) ÷ min × 100, 소수 첫째 자리>(to one decimal place). - From the five runs of step 1, pick the run with the smallest p50 and the run with the largest p50. Write six lines to
/root/lt-regression/02-illusion.txt—fastest_run=<번호>(the number),slowest_run=<번호>(the number),fastest_p50=<밀리초>(milliseconds),slowest_p50=<밀리초>(milliseconds),apparent_change_pct=<(느린 값 - 빠른 값) ÷ 빠른 값 × 100, 소수 첫째 자리>(the slow value minus the fast value, divided by the fast value, times 100, to one decimal place), andsame_code=yes. Write milliseconds to three decimal places. These two were measured on the same version under the same conditions. - Start the same
/opt/lab/lt/lt-regression/svc.pyadditionally on 127.0.0.1:8081 withVERSION=small DELAY_MS=21 JITTER_MS=20 SEED=1234(leave 8080 as it is). Apply exactly the same load as step 1 five times, save as/root/lt-regression/small/run1.csvthroughrun5.csv, and write five lines in the same shape as step 1 to/root/lt-regression/03-small.tsv. Then write one line to/root/lt-regression/03-median.txt,median_of_medians=<다섯 p50 의 중앙값, 소수 셋째 자리>(the median of the five p50s, to three decimal places). The definition of the median is the same as in step 1. - Start
/opt/lab/lt/lt-regression/svc.pyadditionally on 127.0.0.1:8082 withVERSION=big DELAY_MS=30 JITTER_MS=20 SEED=1234, apply the same load five times, and save as/root/lt-regression/big/run1.csvthroughrun5.csv. Write five lines to/root/lt-regression/04-big.tsvand one line,median_of_medians=, to/root/lt-regression/04-median.txt, in the same shape as step 3. - Write four lines to
/root/lt-regression/05-rule.txt—metric=p50,stat=median_of_5,noise_ms=<1단계 다섯 실행의 최대 - 최소, 소수 셋째 자리>(the maximum minus the minimum of the five runs of step 1, to three decimal places), andrule=<이 규칙을 한 문장으로, 40자 이상>(this rule in one sentence, at least 40 characters). Then write three lines to/root/lt-regression/05-verdict.tsv. Each line has three tab-separated columns,<이름> <delta_ms> <judgment>(name, delta in milliseconds, and judgment), and the names are in the orderbase,small,big. delta is the median of that version's five runs minus the median of base's five runs (to three decimal places), and judgment isregressionwhen delta is greater than noise_ms andnoiseotherwise. - Create
/root/lt-regression/gate.sh. When called asbash gate.sh <기준선 디렉터리> <새 측정 디렉터리>(the baseline directory and the new measurement directory), it reads the*.csvof each of the two directories, computes p50 for each run, and gives the median of each group. The output is one line, eitherdelta=<밀리초> noise=<밀리초> REGRESSIONordelta=<밀리초> noise=<밀리초> OK(with milliseconds for each). delta is the median of the new measurement minus the median of the baseline, noise is the maximum of the baseline runs' p50 minus the minimum, and both are to three decimal places. If delta is greater than noise, it exits with code 1, otherwise with 0. The grader runs this script with three sets of inputs it made itself to confirm both passes and failures. - Copy the one with the smallest p50 among the five runs of step 1 to
/root/lt-regression/one-base/run.csvand the one with the largest to/root/lt-regression/one-new/run.csv(both files are measurements of the base version). Run the gate on those two directories, save the output to/root/lt-regression/07-gate-out.txt, and write five lines to/root/lt-regression/07-onerun.txt—noise_ms=<게이트가 낸 값>(the value the gate gave),delta_ms=<게이트가 낸 값>(the value the gate gave),gate_exit=<종료 코드>(the exit code),false_positive=<yes 또는 no>(yes or no), andmin_runs=<기준선을 최소 몇 번 돌려야 한다고 볼 것인가, 2 이상의 정수>(the minimum number of times you think the baseline must be run, an integer of 2 or more). - Run the gate two more times. Save the output of
bash gate.sh base bigto/root/lt-regression/08-gate-big.txtand the output ofbash gate.sh base baseto/root/lt-regression/08-gate-same.txt. Then write six lines to/root/lt-regression/08-decision.txt—big_exit=,same_exit=,rollback=<yes 또는 no, big 판을 되돌릴 것인가>(yes or no: whether to roll back the big version),min_runs=<5 이상의 정수>(an integer of 5 or more),threshold_rule=<문턱을 어떻게 정할 것인가, 40자 이상>(how to set the threshold, at least 40 characters), andwhy=<60자 이상>(at least 60 characters).
Notes
- The working directory is
/root/lt-regression. If it does not exist, create it first. - The load target is
/opt/lab/lt/lt-regression/svc.py. The comment at the top of the file describes the environment variables and how to start the three versions. - Containers cannot be started in this Pod (seccomp). Run the target by starting a Python standard library server directly on
127.0.0.1. To restart a version, kill the previous one withpkill -f 'lt-regression/svc.py'. hey -o csvemits one line per request instead of a summary. The first column is the response time (seconds) and the first line is the header.- Common mistake: giving a different
-nor-cfor each version. If the load conditions differ, the two measurements are not a comparable pair. - Common mistake: fixing the threshold like '5%'. Even with this lab's target, the noise can be larger than that.
- hey (rakyll/hey) · k6 automated performance testing · k6 thresholds · Median absolute deviation · wrk2
Measure five times without changing anything to get the noise
Start /opt/lab/lt/lt-regression/svc.py on 127.0.0.1:8080 with VERSION=base DELAY_MS=20 JITTER_MS=20 SEED=1234. Run hey -n 200 -c 5 -o csv http://127.0.0.1:8080/work five times and save the outputs as /root/lt-regression/base/run1.csv through run5.csv. For each run, compute p50 and write it as five lines to /root/lt-regression/01-noise.tsv — each line has two tab-separated columns, run<번호> <p50 밀리초> (the run number and p50 in milliseconds), with milliseconds to three decimal places. p50 is determined by converting the first column of the csv (response-time, in seconds) to milliseconds, sorting in ascending order, and taking the int(n/2)+1th value (counting from 1). Then write four lines to /root/lt-regression/01-spread.txt: min=, max=, spread_ms=<max - min>, and spread_pct=<(max - min) ÷ min × 100, 소수 첫째 자리> (to one decimal place).
The first line of the csv is the header, so you must skip it. If you sort with gawk's asort, it is done in one line. This target does not burn CPU but just sleeps, so the swing you see here is not machine contention but reproducible swing that the target deliberately puts in. You can think of garbage collection or neighbor load in production as taking that place.
Even the same code diverges this much
From the five runs of step 1, pick the run with the smallest p50 and the run with the largest p50. Write six lines to /root/lt-regression/02-illusion.txt — fastest_run=<번호> (the number), slowest_run=<번호> (the number), fastest_p50=<밀리초> (milliseconds), slowest_p50=<밀리초> (milliseconds), apparent_change_pct=<(느린 값 - 빠른 값) ÷ 빠른 값 × 100, 소수 첫째 자리> (the slow value minus the fast value, divided by the fast value, times 100, to one decimal place), and same_code=yes. Write milliseconds to three decimal places. These two were measured on the same version under the same conditions.
If you sort with sort -k2 -n 01-noise.tsv, both ends come out right away. The percentage that comes out here is 'the amount of change you could claim from a single measurement when nothing has been changed'. If you set a fixed threshold smaller than that, the gate blocks deployments at random.
Measure the version that is 5% slower five times
Start the same /opt/lab/lt/lt-regression/svc.py additionally on 127.0.0.1:8081 with VERSION=small DELAY_MS=21 JITTER_MS=20 SEED=1234 (leave 8080 as it is). Apply exactly the same load as step 1 five times, save as /root/lt-regression/small/run1.csv through run5.csv, and write five lines in the same shape as step 1 to /root/lt-regression/03-small.tsv. Then write one line to /root/lt-regression/03-median.txt, median_of_medians=<다섯 p50 의 중앙값, 소수 셋째 자리> (the median of the five p50s, to three decimal places). The definition of the median is the same as in step 1.
If you change the load conditions, you cannot compare the two versions — keep -n and -c exactly the same as in step 1. The median of five values is the third value after sorting. The reason to use the median of five runs rather than the value of one run is to keep one unlucky run from flipping the verdict.
Measure the version that is 50% slower five times too
Start /opt/lab/lt/lt-regression/svc.py additionally on 127.0.0.1:8082 with VERSION=big DELAY_MS=30 JITTER_MS=20 SEED=1234, apply the same load five times, and save as /root/lt-regression/big/run1.csv through run5.csv. Write five lines to /root/lt-regression/04-big.tsv and one line, median_of_medians=, to /root/lt-regression/04-median.txt, in the same shape as step 3.
Even with three versions up at the same time, they do not interfere with each other — because the targets just sleep, they do not compete for CPU. The difference that comes out in this step is big enough to see with your eyes, but even so, not using 'it is visible' as grounds for the verdict is the point of this lab. The verdict is made by the rule in the next step.
Write the verdict rule in words and apply it to the three versions
Write four lines to /root/lt-regression/05-rule.txt — metric=p50, stat=median_of_5, noise_ms=<1단계 다섯 실행의 최대 - 최소, 소수 셋째 자리> (the maximum minus the minimum of the five runs of step 1, to three decimal places), and rule=<이 규칙을 한 문장으로, 40자 이상> (this rule in one sentence, at least 40 characters). Then write three lines to /root/lt-regression/05-verdict.tsv. Each line has three tab-separated columns, <이름> <delta_ms> <judgment> (name, delta in milliseconds, and judgment), and the names are in the order base, small, big. delta is the median of that version's five runs minus the median of base's five runs (to three decimal places), and judgment is regression when delta is greater than noise_ms and noise otherwise.
The delta of the base line is the difference from itself, so it is 0.000. Do not decide in advance which way small is judged — the noise measured that day decides. That is what differs from a fixed-percentage threshold. The inequality is 'greater than', not 'greater than or equal to'.
Pin the rule down as a CI gate
Create /root/lt-regression/gate.sh. When called as bash gate.sh <기준선 디렉터리> <새 측정 디렉터리> (the baseline directory and the new measurement directory), it reads the *.csv of each of the two directories, computes p50 for each run, and gives the median of each group. The output is one line, either delta=<밀리초> noise=<밀리초> REGRESSION or delta=<밀리초> noise=<밀리초> OK (with milliseconds for each). delta is the median of the new measurement minus the median of the baseline, noise is the maximum of the baseline runs' p50 minus the minimum, and both are to three decimal places. If delta is greater than noise, it exits with code 1, otherwise with 0. The grader runs this script with three sets of inputs it made itself to confirm both passes and failures.
Do not fix in advance the number of csv files in a directory — the grader also runs it with groups that are not five. If there is not a single csv, do not exit with 0; signal it with a different exit code. The case where the new measurement is faster (delta is negative) is not a regression. When comparing values as numbers, use awk's exit code, as in awk 'BEGIN{exit !(d > n)}'.
A single-run comparison makes even the same code a regression
Copy the one with the smallest p50 among the five runs of step 1 to /root/lt-regression/one-base/run.csv and the one with the largest to /root/lt-regression/one-new/run.csv (both files are measurements of the base version). Run the gate on those two directories, save the output to /root/lt-regression/07-gate-out.txt, and write five lines to /root/lt-regression/07-onerun.txt — noise_ms=<게이트가 낸 값> (the value the gate gave), delta_ms=<게이트가 낸 값> (the value the gate gave), gate_exit=<종료 코드> (the exit code), false_positive=<yes 또는 no> (yes or no), and min_runs=<기준선을 최소 몇 번 돌려야 한다고 볼 것인가, 2 이상의 정수> (the minimum number of times you think the baseline must be run, an integer of 2 or more).
If there is only one baseline run, the maximum and the minimum are the same and the noise becomes 0. If the threshold is 0, any difference greater than 0 is a regression. The two measurements reported as a regression in this step were measured on the same version, on the same machine, under the same conditions. In min_runs, write the minimum number of repetitions you will use in your own rule.
Roll back or not — the gate's answer and the person's answer
Run the gate two more times. Save the output of bash gate.sh base big to /root/lt-regression/08-gate-big.txt and the output of bash gate.sh base base to /root/lt-regression/08-gate-same.txt. Then write six lines to /root/lt-regression/08-decision.txt — big_exit=, same_exit=, rollback=<yes 또는 no, big 판을 되돌릴 것인가> (yes or no: whether to roll back the big version), min_runs=<5 이상의 정수> (an integer of 5 or more), threshold_rule=<문턱을 어떻게 정할 것인가, 40자 이상> (how to set the threshold, at least 40 characters), and why=<60자 이상> (at least 60 characters).
Comparing the same data with itself gives a difference of 0, so it must not be a regression under any rule — if the gate blocks even in that case, the implementation is wrong, not the rule. In threshold_rule, write what to use instead of a fixed percentage. In why, write the grounds for deciding to roll back (or not to roll back), connected to the numbers from the previous steps.