Getting a 200 does not mean the work was done
In one line
The numbers of a load test mean something only when the test itself verifies what it measured. A test that looks only at status codes can churn out error pages at 3,000 per second and still say "pass".
Why this matters
The load test results for a new payment API went up on the meeting room screen. 3,012 requests per second, p95 of 8 milliseconds, error rate 0%. Nobody objected, and the deployment went out that week. Then payments stopped 30 minutes after deployment.
The cause was on the test side. The requests the load generator sent had no authentication header, and for requests without authentication the gateway was returning 200 along with {"status":"error","message":"upstream unavailable"}. It had been built that way because an old frontend could not handle 4xx. The load generator looked only at status codes, so it counted all 300,000 as successes.
The numbers looked good for the same reason. The error path touches neither the database nor the queue, so it finishes quickly. So the more errors are mixed in, the better the average and percentiles get. The signal that the test is broken comes not as "it got slower" but as "it got faster". This is the landmine stepped on most often in load testing.
How it works
A load generator knows only three things — how many requests it sent, what the status code of each response was, and how long the response took to come back. The Status code distribution that the hey summary shows is exactly that much. What the body was and whether the server actually did any work cannot be known by the generator.
So you must attach a verification layer to the test yourself. You look at three things together.
| Layer | What it looks at | What happens if you miss it |
|---|---|---|
| Status code | How many responses are not 2xx | A 502 storm is read as "it withstood the load" |
| Body | Whether the response is actually the expected shape | An error inside a 200 is counted as a success |
| Response count | Whether the number of requests the server counted equals the number of responses the generator counted | Requests cut off midway vanish from the statistics |
k6 has built this layer into its standard features under the name checks, and you can hang thresholds on the result to turn it into an exit code. hey has no such feature, so a person has to attach it. Here is one criterion for choosing a tool — is there a place to verify responses?
Attaching verification is not the end. You must also separately ask whether the test conditions are representative of production. Three things in particular often diverge.
Connection reuse. hey reuses TCP connections by default. If you give -disable-keepalive, it opens a new connection for every request. When you send 200 requests with concurrency 10, the former uses 10 connections and the latter 200. If production clients use a connection pool and only the test opens a new connection every time, the test ends up measuring not the service but the TCP handshake and the firewall state table.
Response size. If the test response is 1 KB and the production response is 64 KB, even if requests per second come out similar, bytes per second differ by 64 times. For a service where bandwidth or serialization is the bottleneck, this difference becomes a wrong capacity estimate. You must look at throughput not only in requests per second but also in bytes per second.
Cache keys. If you hit only the same URL repeatedly, everything after the first request is a cache hit. If you set the capacity of production, where the hit rate is 20%, from a test result with a hit rate of 99.5%, the database behind the cache gets deployed without ever having received load in the test.
What it looks like in the field
The most common shape is "test results are better than production". And the cause is almost always that the test did an easier job than production. It skipped authentication, read the same data only, had small responses, or fell into the error path.
There is a case in the opposite direction too. One team tripled its servers because the test result was only a third of production. Later it turned out the load generator had been running with connection reuse turned off. The cost of the tripled servers stayed on the bill.
The third shape is quieter. It is when a test runs fine for months and then one day the results get better. Usually something changed on the target side in the meantime — the authentication method changed so the test token expired, a path was moved so the router is returning a default response, or a feature flag was turned off so the actual computation is skipped. The test is the same but the target started doing a different job, and without a verification layer this change is recorded as "performance improved". So verification is not something you attach once and finish; it has to keep running along with the test.
So a load test report has something that must be written before the numbers. What did this test verify and what did it not verify? If even one thing was not verified, that number is neither an upper nor a lower bound. And that verification must be left not as a person reading the report and confirming, but as the exit code of a script — because people do not doubt good numbers.
What you will do in the next lab
You start a target whose status code is always 200 but whose body is an error one time in three, and first confirm that it looks perfect if you look only at the hey summary. Then you count what actually went out using the access record the server left, and confirm by median that error responses are faster than normal ones. You toggle connection reuse off and on to see the connection count split into 10 and 200, change the response size to 1 KB and 64 KB to compare bytes per second, and with the same key and a rotating key see the cache hit rate split into 99.5% and 0%. Finally you build a checklist script that checks these three things, and the grader runs that script four times with inputs it made itself to confirm both passes and failures.