The report said 500 requests per second; it was actually 48
Goal
You predict the throughput ceiling of a closed-loop generator with Little's law and confirm it by measuring, measure that hey's rate option is a per-worker limit, and build a procedure for telling whether a result that fell short of the target rate is server saturation or a generator limit, plus a report check script.
Why it matters
The number most often wrong in load test reports is not latency but the load itself. A closed loop with concurrency 10 and a 200-millisecond response cannot exceed 50 requests per second no matter how high a target you write. Yet the tool prints pretty percentiles without a warning, and a server that has experienced only a tenth of the real load naturally looks healthy. The definition of the rate option differs between tools, so misreading a single word inflates or shrinks the actual load by the number of workers, and when the generator's own resources, such as the open file count, become the limit, those failures get mixed into the server's error rate. So before reading results, you must first confirm whether the requested load and the load actually applied were the same, and entrust that confirmation not to human memory but to a report template and a check script.
Steps
- Start
/opt/lab/lt/lt-generator-limits/slow.pyon port 8090 with a 0.2-second response. Then call/five times withcurl -s -o /dev/null -w '%{time_total}\n', leave those five lines as they are in/root/lt-generator-limits/01-probe.txt, and write two lines,port=andservice_ms=, to/root/lt-generator-limits/01-service.txt. Writeservice_msas the median of the five lines, an integer in milliseconds. - Create
/root/lt-generator-limits/predict.tsv. It has four lines with no header, and each line has two tab-separated columns,<동시성> <초당 요청 수 상한>(concurrency and the ceiling of requests per second). The concurrencies are, in order, 1, 2, 5, and 10, and for the ceiling write동시성 ÷ 서비스 시간(초)(that is, concurrency ÷ service time in seconds) to two decimal places. The service time is the value you measured in step 1. - Measure twice with
hey -t 60 -o csv. Use-n 30for concurrency 1 and-n 125for concurrency 5, and leave the raw CSVs as they are in/root/lt-generator-limits/c1.csvand/root/lt-generator-limits/c5.csv. Then write two lines to/root/lt-generator-limits/03-measured.tsvas three tab-separated columns,<동시성> <예측 RPS> <측정 RPS>(concurrency, predicted RPS, and measured RPS; to two decimal places). Compute total RPS as요청 수 ÷ (max(offset + response-time) − min(offset))(that is, the number of requests divided by the span from the earliest offset to the latest offset plus response time) — you can get it with only two columns of hey's CSV. - Measure with
hey -n 100 -c 5 -q 2 -t 60 -o csvand leave the raw output in/root/lt-generator-limits/q.csv. Then write five lines,q_flag=,workers=,expected_total_rps=,measured_rps=, andper_worker=, to/root/lt-generator-limits/04-q.txt.expected_total_rpsis the total rate predicted from the-qvalue and the number of workers,measured_rpsis the value recomputed from q.csv (to two decimal places), andper_workerisyesif-qis a per-worker limit andnoif it is a total limit. - Measure with
hey -n 100 -c 2 -q 25 -t 60 -o csvand leave the raw output in/root/lt-generator-limits/shortfall.csv. With-q 25and 2 workers, the requested load is 50 requests per second. Write three lines,requested_rps=,achieved_rps=, andachieved_ratio=, to/root/lt-generator-limits/05-shortfall.txt. For the achievement ratio, writeachieved_rps ÷ requested_rpsto three decimal places. - Write six lines to
/root/lt-generator-limits/verdict.txt.baseline_service_ms=is the median response time ofc1.csv,loaded_service_ms=is the median ofshortfall.csv(an integer in milliseconds),service_ratio=is the ratio of the two (to two decimal places),rps_ratio=is the achievement ratio from step 5,verdict=is the value chosen by the rule below, andreason=is the grounds in at least 60 characters. The rule: ifrps_ratiois 0.8 or more,ok; if not, butservice_ratiois 1.3 or more,server; if neither,generator. - In a shell where the open file limit is lowered to 64, run
hey -n 400 -c 200 -t 5and leave the human-readable output as it is in/root/lt-generator-limits/fdlimit.txt(without-o csv). Then write five lines,nofile=,concurrency=,ok_count=,error_count=, andlimit=, to/root/lt-generator-limits/07-fd.txt. Count the success and error counts from fdlimit.txt, and inlimit=write, asserverorgenerator, which side's limit this failure was. - Create
/root/lt-generator-limits/check-report.sh. It takes a report file path as the first argument and must exit with 0 if the three linesrequested_rps=,achieved_rps=, andachieved_ratio=are all present with numbers, and with a non-zero value if even one is missing. Then write a report,/root/lt-generator-limits/report.md, that passes that check, and its three values must be the same as those measured in step 5.
Notes
- The working directory is
/root/lt-generator-limits. If it does not exist, create it first. - The load target is just
/opt/lab/lt/lt-generator-limits/slow.py. A Pod cannot start containers, so it is run by starting a Python standard library HTTP server directly on 127.0.0.1. - In hey's CSV, the first line is the header, the first column is
response-time, and the last column isoffset. You get total RPS with요청 수 ÷ (max(offset + response-time) − min(offset))(that is, the number of requests divided by the span from the earliest offset to the latest offset plus response time). - Common mistake: reading the rate option as a total target rate. Step 4 is that check.
- Common mistake: writing a result that fell short of the target as 'the server's limit' as it is. If the server-side processing time is unchanged, the limit is on the generator side.
- hey — definitions of the options -n, -c, and -q · k6 — open vs closed model · k6 — scenarios and arrival rates · vegeta — fixed-rate attacks · wrk2 — target rate and correction
Start a load target with a constant service time
Start /opt/lab/lt/lt-generator-limits/slow.py on port 8090 with a 0.2-second response. Then call / five times with curl -s -o /dev/null -w '%{time_total}\n', leave those five lines as they are in /root/lt-generator-limits/01-probe.txt, and write two lines, port= and service_ms=, to /root/lt-generator-limits/01-service.txt. Write service_ms as the median of the five lines, an integer in milliseconds.
This target has no lock, so it handles as many requests as arrive at the same time together — in this lab the limit is not on the server but on the side that generates the load. To run it in the background, use nohup ... & and wait about 1 second until it comes up. The median is the middle line after sort -n.
Predict the ceiling first with Little's law
Create /root/lt-generator-limits/predict.tsv. It has four lines with no header, and each line has two tab-separated columns, <동시성> <초당 요청 수 상한> (concurrency and the ceiling of requests per second). The concurrencies are, in order, 1, 2, 5, and 10, and for the ceiling write 동시성 ÷ 서비스 시간(초) (that is, concurrency ÷ service time in seconds) to two decimal places. The service time is the value you measured in step 1.
In a closed loop, one worker does not send the next request while it waits for a response. So the maximum rate C workers can produce is C ÷ response time (Little's law). This ceiling is independent of server performance — no matter how fast the server is, requests do not go out while waiting.
Actually measure whether the prediction is right
Measure twice with hey -t 60 -o csv. Use -n 30 for concurrency 1 and -n 125 for concurrency 5, and leave the raw CSVs as they are in /root/lt-generator-limits/c1.csv and /root/lt-generator-limits/c5.csv. Then write two lines to /root/lt-generator-limits/03-measured.tsv as three tab-separated columns, <동시성> <예측 RPS> <측정 RPS> (concurrency, predicted RPS, and measured RPS; to two decimal places). Compute total RPS as 요청 수 ÷ (max(offset + response-time) − min(offset)) (that is, the number of requests divided by the span from the earliest offset to the latest offset plus response time) — you can get it with only two columns of hey's CSV.
In hey's CSV, the first line is the header, the first column is response-time, and the last column is offset (the time since the start of the test when that request was sent). It is normal for the measured value not to exceed the predicted value — if it did, you measured the service time wrongly or the target cut off responses early.
-q is a per-worker limit, not a total target rate
Measure with hey -n 100 -c 5 -q 2 -t 60 -o csv and leave the raw output in /root/lt-generator-limits/q.csv. Then write five lines, q_flag=, workers=, expected_total_rps=, measured_rps=, and per_worker=, to /root/lt-generator-limits/04-q.txt. expected_total_rps is the total rate predicted from the -q value and the number of workers, measured_rps is the value recomputed from q.csv (to two decimal places), and per_worker is yes if -q is a per-worker limit and no if it is a total limit.
If -q were a total limit, the total rate should be near 2. See what actually comes out and judge. Also check how hey's documentation describes this option — a difference of one word turns the whole test into a different experiment. You must set it lower than the ceiling of concurrency 5 (25) for the limit to actually take effect.
Create a test where less than half of the target was applied
Measure with hey -n 100 -c 2 -q 25 -t 60 -o csv and leave the raw output in /root/lt-generator-limits/shortfall.csv. With -q 25 and 2 workers, the requested load is 50 requests per second. Write three lines, requested_rps=, achieved_rps=, and achieved_ratio=, to /root/lt-generator-limits/05-shortfall.txt. For the achievement ratio, write achieved_rps ÷ requested_rps to three decimal places.
With concurrency 2 and a 0.2-second response, the Little's law ceiling is 10 requests per second. No matter how high you set the rate limit, it cannot rise above that — the rate option means 'do not send faster than this', not 'send this many'.
Is it server saturation or a generator limit
Write six lines to /root/lt-generator-limits/verdict.txt. baseline_service_ms= is the median response time of c1.csv, loaded_service_ms= is the median of shortfall.csv (an integer in milliseconds), service_ratio= is the ratio of the two (to two decimal places), rps_ratio= is the achievement ratio from step 5, verdict= is the value chosen by the rule below, and reason= is the grounds in at least 60 characters. The rule: if rps_ratio is 0.8 or more, ok; if not, but service_ratio is 1.3 or more, server; if neither, generator.
When a server saturates, processing time grows. If the processing time is unchanged and only the total rate falls short of the target, the server is still idle and the limit is on the side that generates the load. If you do not make this distinction, you end up attaching more instances to a perfectly healthy server.
The moment the generator itself becomes the limit
In a shell where the open file limit is lowered to 64, run hey -n 400 -c 200 -t 5 and leave the human-readable output as it is in /root/lt-generator-limits/fdlimit.txt (without -o csv). Then write five lines, nofile=, concurrency=, ok_count=, error_count=, and limit=, to /root/lt-generator-limits/07-fd.txt. Count the success and error counts from fdlimit.txt, and in limit= write, as server or generator, which side's limit this failure was.
ulimit -n 64 applies only to that shell and its child processes — bundle it into one line like bash -c 'ulimit -n 64; hey ...'. The output of hey has two sections, Status code distribution and Error distribution. If you read the error messages, you can tell who raised this failure.
Make the report template ask for itself
Create /root/lt-generator-limits/check-report.sh. It takes a report file path as the first argument and must exit with 0 if the three lines requested_rps=, achieved_rps=, and achieved_ratio= are all present with numbers, and with a non-zero value if even one is missing. Then write a report, /root/lt-generator-limits/report.md, that passes that check, and its three values must be the same as those measured in step 5.
The grader feeds this script a report with items deliberately removed and sees whether it fails — a script that always exits with 0 does not pass. One small check like this makes 'we applied 500 requests per second' prove itself. Also write the test conditions in the report.