TT Lab
Get started
Learn Learning paths Courses

Load Testing

An average of 100 requests per second does not describe the load

Continue in TT Lab

In one line

"25 requests per second" explains only half of the load. Even with the same average, the queue and latency differ depending on whether arrivals are even or bunched, and in a closed loop you cannot even set the request rate.

Why this matters

Capacity estimation usually starts like this — the peak is 25 requests per second and each takes 40 milliseconds, so the workers needed are 25 times 0.04, which is 1. To be generous, we put 4. The calculation is right. Yet when deployed, the queue shoots up to dozens and p95 jumps tenfold.

What is wrong is not the arithmetic but the assumption. That calculation assumes requests arrive one every 40 milliseconds, evenly. Real traffic does not arrive that way. A service used by people gets crowded by a single notification, and a service used by machines gets crowded when cron wakes up on the hour. The average is the same but the number of simultaneous requests at a moment is different, and a queue reacts not to the average but to the moment.

There is a misunderstanding in the opposite direction too. It is the case where you enter "100 concurrent users" into a load test tool and write down in advance how many requests per second will come out. In a closed loop, the request rate is not an input but a result. When the server slows down, users send their next request later, so the request rate falls by itself, and so the test draws a picture that "it holds up well" no matter how broken the server is.

How it works

An arrival process is the shape of "when requests come". Knowing three is enough.

Shape Interval Coefficient of variation (CV) Where you see it
Constant interval Always 1 / request rate 0 The default of synthetic load tests
Exponential (Poisson) Exponential distribution with mean 1 / request rate 1 Many mutually independent users
Bunched (burst) Several at once, then a long pause Much greater than 1 Cron, retry storms, notification sends

The coefficient of variation is the standard deviation of the intervals divided by the mean. Even with the same average request rate, the larger this value, the longer the wait. The direction the queueing theory approximation (Kingman) points is the same — wait time is proportional not only to utilization but also to the variability of arrival and of service. So the sentence "utilization is 25%, so it is safe" means nothing unless it states the arrival shape.

For the closed loop there is a different law. If N users each send one request and then pause for think time Z, and the response time is R, the effective request rate X is N / (R + Z). This is Little's law applied to one user's round trip, meaning that once you fix the number of users and the think time, the request rate is decided by the server. If you leave out R and predict with N / Z, it always comes out higher than reality, and the slower the server, the larger that error.

The k6 documentation names these two by executor. The arrival-rate executor, which takes a request rate as input, is the open model, and the executor that takes a number of virtual users is the closed model. Which one to use is decided not by taste but by what you are trying to measure.

What it looks like in the field

The accident seen most often is a retry storm. Requests that normally arrive evenly hit a timeout once, and the clients retry at the same moment, so arrivals turn into bunches. The average request rate grew by only a few percent, but the queue became ten times longer, and that wait produces timeouts again. This is why retries without jitter are prohibited.

Batches that wake up on the hour are the same. If the cron expressions are all 0 * * * *, even if the daily average load is low, the number of simultaneous requests on the hour every hour far exceeds the number of workers. The capacity of such a service must be set not by the average but by the size of the bunch.

And there is one more accident on the test tool side. If you test with a number of virtual users a service that should be tested with an arrival-rate executor, the load drops along with it when the server slows down and the saturation point is not visible. The report says "it held up to so many requests per second," but that number is not the server's limit but the limit of the test configuration.

So when writing down load in a document, two things must be written next to the average. One is the shape of the arrival — even, mutually independent, or bunched. The other, if it is bunched, is how many at a time. Without these two, "25 requests per second" is a number you cannot use to decide the number of workers. The way to extract this value from observation is also simple. If you group the times in the access log into short windows (for example 100 milliseconds) and count them, even traffic gives similar counts per window, while bunched traffic gives mostly 0 with only a few windows spiking high. The size of those spiking windows is exactly the number to put into the capacity calculation.

The remaining question is how much margin to leave. The right answer depends on what the service promised. If it is a request a user is waiting on, you must have enough workers to take the bunch as it is, and if it is background work, you can keep a long queue and digest it slowly. What matters is making that choice on a measured bunch size, not looking only at the average.

What you will do in the next lab

You start a server with a fixed number of workers and create and send three arrivals with the same average request rate yourself: constant interval, exponential, and bunched. After measuring the coefficient of variation of the intervals of the three plans and confirming that the average is the same and only the shape differs, you compare the queue length and response time distribution recorded by the server. Next you build a closed-loop generator and compare, against the law and against measurement, how think time decides the effective request rate, and finally you show in numbers how many times off a capacity estimate that assumes even arrivals is under real bunched arrivals, and then decide this service's arrival model and number of workers together with the evidence.