TT Lab
Get started
Learn Learning paths Courses

Load Testing

Load the generator never produced is not the server's limit

Continue in TT Lab

In one line

Load that the load generator could not produce is not the server's limit. Before reading test results, first confirm whether the requested load and the load actually applied were the same.

Why this matters

The capacity review report of an order API said this. "We applied 500 requests per second for 1 minute, p99 was 210 milliseconds, and there were 0 errors." Yet the server's request count metric for the same period was flat at 48 requests per second. 500 had never been applied.

The cause was simple. The generator was run at concurrency 10 and the response took 200 milliseconds. The maximum throughput this combination can produce in a closed loop is 10 ÷ 0.2 = 50 requests per second. 500 was just a wish written on the command line, a number that physically could not come out. Yet the tool, without any warning, prettily printed "p99 210 milliseconds".

Such a result is dangerous because it creates a wrong belief about the server. A server that has experienced only a tenth of the real load naturally looks healthy. If you size capacity on that number, on launch day it collapses by exactly a tenfold gap.

How it works

The throughput ceiling of a closed-loop generator comes straight from Little's law. Since L = λ × W holds among the number of requests waiting L, the arrival rate λ, and the time spent W, if you fix the concurrency at C, λ cannot exceed C ÷ 응답시간 (that is, C divided by the response time).

Concurrency Ceiling when the response is 200 ms
1 5 per second
5 25 per second
10 50 per second

What matters here is that this ceiling is independent of the server. No matter how fast the server is, requests do not go out while the generator is waiting. So if you want to reach a target rate, you must set the concurrency at or above the target rate × response time.

The second trap is the meaning of the rate option. The -q of hey is documented as "Rate limit, in queries per second (QPS) per worker" — a limit applied per worker. -c 5 -q 2 is not 2 per second but 10 per second. Conversely, if you believe you entered a total target rate, the actual load is inflated by the number of workers. The definition differs between tools, so misreading a single option turns the whole test into a different experiment.

The third is the generator's own resources. One request needs one socket, so if the concurrency exceeds the open file limit, that many fail quietly. The number of processes, CPU, and the network bandwidth of the host the generator runs on also become ceilings in the same way. If the generator runs on a machine smaller than the target, that test is a test of the generator.

So there is an order when designing a test. First decide the target rate you want to apply, multiply that rate by the response time to get the required concurrency. For 500 requests per second with a 200-millisecond response, at least 100 workers are needed. Then confirm that this concurrency fits within the generator-side limits — open file count, ephemeral port range, and memory, in that order. If you do these two calculations first, there is less to investigate later about "why the target was not applied".

What it looks like in the field

The most common signal is a result where "I doubled the concurrency and the total throughput also exactly doubled and the response time stayed the same". If the server were saturated, throughput would flatten and response time would rise. If it is neither, the server is still idle and the limit is on the generator side.

So the judgment procedure needs only two numbers. If the server-side processing time is unchanged and only the total rate falls short of the target, it is a generator limit. Conversely, if processing time grows while the rate flattens, that is the real saturation point. If you do not make this distinction, you end up attaching more instances to a perfectly healthy server.

The error distribution is also a clue. "connection refused" or "too many open files" are often errors raised by the generator, not the server. The file descriptor limit in particular appears suddenly the moment you raise concurrency a lot, and the tool just counts it as failed requests and mixes it into the error rate. Then the conclusion "the server failed 5%" comes out.

Another common picture is a test where the generator and the target run on the same machine. This lab's Pod is exactly that situation. When the two sides share the same CPU, the more you raise the load, the smaller the generator's share becomes, and so you cannot tell whether the target is slowing down or the generator is. In a real test, the generator should be placed on a different machine, and if that is not possible, at least this fact must be written in the report.

Finally, the cheapest way to prevent all of this is the report template. If you create a field that puts the requested load and the load actually applied side by side, anyone will notice when the two differ. Without the field, nobody asks.

What you will do in the next lab

You start a Python server whose response takes 200 milliseconds, first predict the throughput ceiling per concurrency with Little's law, and then actually measure with hey to check. You confirm by measurement that -q is a per-worker limit, deliberately create a test that falls far short of the target rate, and establish a procedure for telling whether it is server saturation or a generator limit. You also create yourself the situation in which the generator itself becomes the limit by lowering the open file limit, and finally build and leave a check script that verifies whether the requested load and the load actually applied are written together.