The Load Generator Must Cooperate With Backpressure
In one line
When load test results and production metrics differ, you should usually suspect the load test. Real users do not delay the next request just because the previous one is slow.
Why this matters
Suppose you set the load generator to 1,000 requests per second and one response takes 2 seconds. A synchronous (closed) generator does not send the 2,000 requests it should have sent during those 2 seconds. The samples from the period when the system was slowest vanish entirely, and the load generator ends up 'cooperating' with the backpressure it created itself. This is the problem Gil Tene called coordinated omission.
There are three symptoms. Raising the load barely moves p99. The reported throughput is lower than the configured target but the latency distribution looks as if it was measured at the target. And production has a much worse tail than the load test.
The acceptance check to apply to every run is one line. Does the reported actual throughput match the configured target throughput? If it does not, the latency distribution of that run cannot be trusted.
How it works
This is the difference between the closed model and the open model. Closed fixes the concurrency (always N in progress). Open fixes the arrival rate (λ requests arrive per second regardless of whether responses have come back). Real users are closer to open.
Little's law L = λ × W connects the two models. The average number of requests in the system is the arrival rate times the average time spent. At RPS 1,000 and latency 100ms, 100 are in progress at the same time. Conversely, if in a closed run with concurrency fixed at 10, RPS × average latency is not near 10, it means the load generator did not actually produce the target load.
It is used in exactly the same way for connection pool sizing. If one instance handles 500 requests per second at an average of 80ms, then 500 × 0.08 = 40, and giving a 1.5x margin to account for tail latency, the recommended pool size is 60. You must not stop here. You must check again at the fleet level. 40 instances × a pool of 60 = at most 2,400 connections, and if Postgres max_connections is 200, a connection storm right after deployment causes an outage.
What it looks like in the field
Bad benchmarks have things in common. They have no warm-up so the JIT has not heated up, the CPU frequency is not pinned, OS noise is not accounted for, they were measured only once, and outliers were not handled.
A good procedure is simple. 10 warm-up runs, 100 or more measured runs, removal of anything beyond 2σ, reporting p50, p95, and p99 together, and recording the environment. Knowing measured cold start values also helps in deciding the warm-up length. Python/Node 150–500ms, Go/Rust 50–100ms, JVM 1000–3000ms, GraalVM Native 50–100ms.
Test types are also divided by purpose. Load (normal operation at expected traffic), stress (gradual increase to find the breaking point), spike (reaction to a sudden surge), and endurance, or soak (long duration to detect memory leaks).
How far to climb the ladder
The mistakes people make most often on a ladder where you measure while raising concurrency are stopping too early and dragging on too long. If you stop while throughput is still rising, you do not find the limit, and if you keep raising after it has already bent, the test itself breaks the system and all later measurements become unusable.
There are three criteria for deciding where to stop.
- Throughput no longer rises. If you doubled the concurrency and throughput grew by less than 10%, it is already saturated. That point is called the knee.
- Errors start to appear. Once the error rate rises, the latency numbers after that are meaningless. Failed requests usually finish quickly, so there is even the illusion that latency looks better.
- It does not come back. If you lowered the load and the latency does not return to its original level, something is broken. Queues have built up, connections are leaking, or memory has not been reclaimed. At that point, you must stop the test and look at the cause, not keep raising the ladder.
The third is in fact the most valuable observation. A healthy system's latency comes down when you lower the load. If it does not, it means the system cannot recover by itself when it fails to withstand a peak, and that is a far more important fact than the throughput number. So after climbing the ladder, you must measure once more on the way down. If the values differ at the same concurrency going up and coming down, that difference is itself the flaw in resilience.
One more thing. Each step must be long enough. With a 30-second step, you observe neither autoscaling, nor cache warming, nor queues building up. You must hold at least a few minutes for steady-state values to appear, and it is better to mark the values before that separately as a transient period. In a system with autoscaling in particular, the time until a new instance starts and becomes ready is entirely transient, so with steps shorter than that time, you cannot even tell whether scaling actually helps.
What you will do in the next lab
You plan a concurrency ladder of 1, 2, 5, 10, and 20 and run it automatically with a script. You make the results into a table, find the knee, where throughput no longer grows, by a rule, verify with Little's law that the run itself was valid, and then run a separate open run with a fixed arrival rate and compare the actual throughput against the target.