The server stalled for two seconds and p99 was 60 milliseconds
Goal
You measure the same server once with a closed loop and once with an open model to see for yourself how far p99 diverges, implement by hand an HdrHistogram-style correction on the closed-loop samples, confirm the SLA verdict flipping before and after the correction, and leave test design rules in a file.
Why it matters
A closed-loop load generator sends the next request only after receiving the previous request's response. So while the server is stalled it sends no requests at all, and that period barely remains in the latency distribution. Real users do not hold back their requests depending on the server's condition, so this measurement is always wrong only in the direction of looking better than reality. It is not an error that wobbles at random but one that leans in only one direction, so nobody suspects it and it passes the launch review. There are two ways to fix it — measure with an open model that pins the send time from the start, or fill in the samples you have already measured using the expected interval. You must do one of the two, and the report must record which one it was.
Steps
- Start
/opt/lab/lt/lt-coordinated-omission/target.pyon port 8080 (normal response 0.05 seconds, stall 2.0 seconds, stalling on the 150th request). Then call/six times withcurl -s -o /dev/null -w '%{time_total}\n', leave those six lines as they are in/root/lt-coordinated-omission/01-probe.txt, and write four lines,port=,base_ms=,stall_ms=, andstall_at=, to/root/lt-coordinated-omission/01-target.txt.base_msis the median of the six lines as an integer in milliseconds, andstall_msandstall_atare values read from the server source. - After resetting the request counter (
/reset), measure withhey -n 200 -c 1 -q 12.5 -t 60 -o csvand leave the raw CSV as it is in/root/lt-coordinated-omission/closed.csv. Using only theresponse-timecolumn of that CSV, write four lines,samples=,p50_ms=,p99_ms=, andmax_ms=, to/root/lt-coordinated-omission/02-closed.txt. Round to an integer in milliseconds. Compute percentiles as the value at the position인덱스 = int(백분위 ÷ 100 × 표본수)(that is, index = int(percentile ÷ 100 × number of samples), counting from 0 and using the last one if it exceeds the number of samples − 1) after sorting in ascending order. This is the same method hey uses. - Create 200 lines in
/root/lt-coordinated-omission/schedule.txt. One per line, the time at which you intended to send request i, in seconds to three decimal places. The first line is0.000and the interval is a constant 0.080 seconds. This file becomes the basis for corrected latency in the next step. - Write
/root/lt-coordinated-omission/openloop.pyyourself. It sends one request at each time inschedule.txtbut does not wait for the previous request. Leave the results in/root/lt-coordinated-omission/open.csvwith the headeri,intended,sent,received,codeand 200 lines (times are seconds from the start of measurement), and write four lines,samples=,p50_ms=,p99_ms=, andmax_ms=, of the corrected latency (received − intended) to/root/lt-coordinated-omission/04-open.txt. - Write
/root/lt-coordinated-omission/hdrfix.pyto fill in the latency samples ofclosed.csvwith an expected interval of 0.080 seconds. If one sample's value is larger than the expected interval, keep subtracting the expected interval from that value and create more samples while it is greater than 0 (what HdrHistogram'srecordValueWithExpectedIntervaldoes). Leave the filled-in result, one per line in seconds, in/root/lt-coordinated-omission/closed-corrected.txt, and write five lines,expected_interval_ms=,samples=,p50_ms=,p99_ms=, andmax_ms=, to/root/lt-coordinated-omission/05-hdr.txt. - Create
/root/lt-coordinated-omission/compare.tsv. It has three lines with no header, and each line has four tab-separated columns,<이름> <p50_ms> <p99_ms> <max_ms>(name, p50, p99, and max in milliseconds). The names are, in order,closed(step 2),hdr(step 5), andopen(step 4), and you copy the numbers as the integer milliseconds you wrote in the previous steps. - Write six lines to
/root/lt-coordinated-omission/sla.txt. Forsla_p99_ms=, an integer threshold between 100 and 1000; forclosed_p99_ms=andcorrected_p99_ms=, the p99 of the closed loop (step 2) and of the open-model correction (step 4), respectively; forclosed_verdict=andcorrected_verdict=,passif that p99 is at or below the threshold andfailif it is above; and forflipped=,yesif the two verdicts differ andnoif they are the same. - Write two lines to
/root/lt-coordinated-omission/08-summary.txt,worst_case_ms=(the maximum of the open-model corrected latency) andomitted_requests=(the number of requests the closed loop could not send while stalled = the quotient of the largest latency sample divided by the expected interval of 0.080 seconds). Then write four rules in the form- <열쇠>: <설명>(a dash, a key, a colon, and a description) to/root/lt-coordinated-omission/rules.md. The keys are, in order,rate,correct,report, andgate, and each description must be at least 40 characters.
Notes
- The working directory is
/root/lt-coordinated-omission. If it does not exist, create it first. - The load target and traffic data are in
/opt/lab/lt/lt-coordinated-omission/— just one file,target.py. A Pod cannot start containers, so the load target is run by starting a Python standard library HTTP server directly on 127.0.0.1. - Before starting measurement, reset the request counter with
curl http://127.0.0.1:8080/reset. If you do not, the position of the stalling request changes. - Common mistake: guessing that because the maximum is large, p99 will also be large. With 200 samples, one bad sample is not captured in p99. Always look at the two numbers together.
- Common mistake: writing the open-model script so that it waits for the previous response. If
sentinopen.csvdrifts away fromintended, it is not an open model but a slow closed loop. - How to use hey and its options · wrk2 — coordinated omission correction · HdrHistogram · k6 — open vs closed model · vegeta — fixed-rate attacks
Start a load target that stalls periodically
Start /opt/lab/lt/lt-coordinated-omission/target.py on port 8080 (normal response 0.05 seconds, stall 2.0 seconds, stalling on the 150th request). Then call / six times with curl -s -o /dev/null -w '%{time_total}\n', leave those six lines as they are in /root/lt-coordinated-omission/01-probe.txt, and write four lines, port=, base_ms=, stall_ms=, and stall_at=, to /root/lt-coordinated-omission/01-target.txt. base_ms is the median of the six lines as an integer in milliseconds, and stall_ms and stall_at are values read from the server source.
To start it in the background, use nohup python3 ... > server.log 2>&1 &. Wait 1–2 seconds until it comes up. If 8080 is already open, you cannot start it twice — if curl http://127.0.0.1:8080/reset returns reset, it is already running. For the median, look at the middle line after sort -n.
Measure with a closed loop and write down the raw percentiles
After resetting the request counter (/reset), measure with hey -n 200 -c 1 -q 12.5 -t 60 -o csv and leave the raw CSV as it is in /root/lt-coordinated-omission/closed.csv. Using only the response-time column of that CSV, write four lines, samples=, p50_ms=, p99_ms=, and max_ms=, to /root/lt-coordinated-omission/02-closed.txt. Round to an integer in milliseconds. Compute percentiles as the value at the position 인덱스 = int(백분위 ÷ 100 × 표본수) (that is, index = int(percentile ÷ 100 × number of samples), counting from 0 and using the last one if it exceeds the number of samples − 1) after sorting in ascending order. This is the same method hey uses.
-c 1 means one worker, and -q 12.5 means the limit is 12.5 requests per second per worker (an interval of 80 milliseconds). If you give -o csv, you get one line per request instead of a summary, and the first line is the header. See by how many times p99 and the maximum differ — that gap is the subject of this lab.
Pin down the intended send times in advance
Create 200 lines in /root/lt-coordinated-omission/schedule.txt. One per line, the time at which you intended to send request i, in seconds to three decimal places. The first line is 0.000 and the interval is a constant 0.080 seconds. This file becomes the basis for corrected latency in the next step.
A closed loop has no such thing as an 'intended time' at all — it sends when the previous response arrives. An open model decides that time first and keeps it regardless of the server's condition. You can make it with one line of seq or awk. The last line is (200 − 1) × 0.08.
Measure again with an open model and get the corrected latency
Write /root/lt-coordinated-omission/openloop.py yourself. It sends one request at each time in schedule.txt but does not wait for the previous request. Leave the results in /root/lt-coordinated-omission/open.csv with the header i,intended,sent,received,code and 200 lines (times are seconds from the start of measurement), and write four lines, samples=, p50_ms=, p99_ms=, and max_ms=, of the corrected latency (received − intended) to /root/lt-coordinated-omission/04-open.txt.
Create one thread per request, and have that thread time.sleep until its own time and then send. urllib.request.urlopen is enough. Use only the standard library (pip install is not possible). It is an open model only if sent and intended are almost equal — if they drift apart, the generator could not keep up.
Implement HdrHistogram's correction by hand
Write /root/lt-coordinated-omission/hdrfix.py to fill in the latency samples of closed.csv with an expected interval of 0.080 seconds. If one sample's value is larger than the expected interval, keep subtracting the expected interval from that value and create more samples while it is greater than 0 (what HdrHistogram's recordValueWithExpectedInterval does). Leave the filled-in result, one per line in seconds, in /root/lt-coordinated-omission/closed-corrected.txt, and write five lines, expected_interval_ms=, samples=, p50_ms=, p99_ms=, and max_ms=, to /root/lt-coordinated-omission/05-hdr.txt.
A single sample with a latency of 2.0 seconds creates several samples at an expected interval of 0.08 seconds: 1.92, 1.84, and so on. Leave the original sample as it is and add the ones you created. The expected interval is the interval you set with -q in step 2 (1 ÷ 12.5 = 0.08 seconds). The percentile calculation method is the same as in the previous step.
Put the three distributions in one table
Create /root/lt-coordinated-omission/compare.tsv. It has three lines with no header, and each line has four tab-separated columns, <이름> <p50_ms> <p99_ms> <max_ms> (name, p50, p99, and max in milliseconds). The names are, in order, closed (step 2), hdr (step 5), and open (step 4), and you copy the numbers as the integer milliseconds you wrote in the previous steps.
The key point is that the p50 of the three lines is almost the same and only p99 diverges. That p50 is the same means the server was running well normally, and that p99 diverges means the samples of the bad period are present on one side and absent on the other. The grader recomputes the values from the three raw files and checks them.
The same test, a flipping SLA verdict
Write six lines to /root/lt-coordinated-omission/sla.txt. For sla_p99_ms=, an integer threshold between 100 and 1000; for closed_p99_ms= and corrected_p99_ms=, the p99 of the closed loop (step 2) and of the open-model correction (step 4), respectively; for closed_verdict= and corrected_verdict=, pass if that p99 is at or below the threshold and fail if it is above; and for flipped=, yes if the two verdicts differ and no if they are the same.
It does not matter where you set the threshold, but if you choose a place where the two verdicts split, the point of this step shows. If the verdict flipped, what needs fixing is not the threshold but 'which number the promise was made in'. If you do not write whether correction was applied in the report, nobody can tell the difference half a year later.
Test design rules so you are never fooled again
Write two lines to /root/lt-coordinated-omission/08-summary.txt, worst_case_ms= (the maximum of the open-model corrected latency) and omitted_requests= (the number of requests the closed loop could not send while stalled = the quotient of the largest latency sample divided by the expected interval of 0.080 seconds). Then write four rules in the form - <열쇠>: <설명> (a dash, a key, a colon, and a description) to /root/lt-coordinated-omission/rules.md. The keys are, in order, rate, correct, report, and gate, and each description must be at least 40 characters.
The questions the four rules must answer are these — which model to apply load with (rate), how to correct samples already measured (correct), what to write in the report along with them (report), and which number to use to judge a pass (gate). Write each line in your own words, but base it on the numbers you saw in the previous steps.