TT Lab
Get started
Learn Learning paths Courses

Load Testing

The Average Cannot Detect the Tail Getting Worse

Continue in TT Lab

In one line

Looking at the same log, a team that looks only at the average decides to roll back, a team that looks only at the median presents it as a success story, and a team that looks only at p99 opens an incident. All three looked at the same data.

Why this matters

Consider a service where 9,900 of 10,000 requests take 50ms and 100 take 3,000ms. The average is 79.5ms. Not a single request actually got a response in 79.5ms. The average describes a user who does not exist. p50, p95, and p99 are all 50ms, and only p99.9 is 3,000ms.

Sensitivity is worse. Even if the slow 100 get twice as bad, from 3,000 to 6,000ms, the average moves only from 79.5 to 109.5ms. The average hardly detects the tail getting worse.

If you look at real numbers from before and after a rollout, it becomes clear why the conclusions split. Average 112 → 122ms (+9%), p50 99 → 54ms (−46%), p95 224 → 454ms (+103%), p99 309 → 678ms (+119%). Looking only at the average, it is a slight degradation; looking only at the median, it is a big improvement; looking at p95/p99, it is a serious regression.

How it works

If you understand p99 as "the 1% of unlucky users", you set priorities wrongly. Tails are amplified. When you make n requests, the probability of hitting the p99 range at least once is 1 − 0.99^n.

Number of calls Probability of hitting p99
1 1.0%
5 4.9%
20 18.2%
50 39.5%
200 86.6%

If one dashboard screen calls 20 APIs, the probability of hitting p99 latency each time you open the screen is 18%. For an active user who makes 200 requests a day, 87% experience the worst range at least once a day. p99 is not the fate of a few unlucky users but the everyday experience of almost every user. So if one screen calls 20 APIs, the target for a single service must be 99.9%, not 99%, for the screen level to come out at 98%.

Percentiles cannot be combined either. The p99 of server A (all 9,900 requests at 50ms) is 50, and the p99 of server B (all 100 requests at 3,000ms) is 3,000. The simple average is 1,525ms, the request-count-weighted average is 79.5ms, and the true combined p99 is 3,000ms. All three numbers are different, and the first two are meaningless.

What it looks like in the field

The most common flaw in load test reports is that there are no conditions. Only RPS and p95 are written, with no concurrency, no total number of requests, and no time of measurement. Such numbers cannot be compared next month, so they are not a baseline.

The second is measuring only once. If you run the same conditions three times, you see the spread, and a difference smaller than that spread cannot be called an improvement.

How to store percentiles correctly

The problem that p99 cannot be combined mostly disappears if you change how you store it. The key is to store the distribution, not the percentile value. A histogram holds counts per bucket, such as "how many in 0–10ms, how many in 10–25ms", so if you simply add the buckets of several instances, you get the whole distribution. If you compute the percentile from that combined distribution, that is the true value. The same goes for the time axis: if you add up 5-minute histograms for a day, you get that day's p99.

In exchange, histograms have an error of their own. You can only know which bucket the percentile falls in, and the exact position within it is estimated by interpolation. So placing bucket boundaries densely near the target determines the accuracy. If the target is 300ms and the boundaries go from 100ms straight to 1 second, every value in between is lumped into one bucket and you cannot judge p99. Conversely, there is no harm in leaving the range above 10 seconds, which nobody looks at, sparse.

And the maximum cannot be obtained from a histogram. The last bucket is open, like "1 second or more", so you cannot tell whether there is a 1.1-second or a 30-second value inside it. In investigations where how slow the slowest request was matters, you must either export the maximum separately as a metric or find the slow requests themselves in logs and traces.

What you measure comes before the numbers

Even if you read percentiles correctly, a number measured under wrong conditions in the first place is useless. The places where a load test diverges from reality are usually well known.

No warm-up. JIT compilation, filling the connection pool, cache warming, and the time until autoscaling reacts. The first 30 seconds after start are not a steady state, so if you include that period in the aggregate, p99 comes out much worse than reality. Conversely, an excessively long warm-up is also a problem. Real users send requests right after a deployment too, so you must measure that period separately and know its performance.

The data is too clean. If every request looks up the same user and the same product, everything after the first request comes from the cache. The throughput number is beautiful but a value that does not exist in reality. Conversely, if you use completely random keys every time, the cache hit rate becomes 0 and it is overly pessimistic. Real access distributions are usually concentrated on a few items, so you must imitate that skew.

Hitting only one place. If you concentrate load on one endpoint, you see only that path's bottleneck. In reality, several features share the same database and the same thread pool, so something that was fine alone collapses when run together. Scenarios must be mixed at the real traffic ratios.

The load tool tires first. If the client-side CPU saturates or ports are exhausted, the limit is the tool, not the server. The latency increase that appears then is not a property of the server. You can avoid this illusion only by also recording the resources of the machine the load tool runs on.

Finally, it is good to fix the format for recording results, too. Load, duration, total number of requests, error rate, p50/p95/p99, and the server's CPU, memory, and connection count at the time. Without this combination, you cannot compare even if you run the same test next month, and a number that cannot be compared is not a baseline but merely a record of that day.

What you will do in the next lab

You start an nginx container as the load target and run a smoke run, a warm-up, and the main measurement separately with hey. You extract RPS, p50/p95/p99, and the error rate yourself from the tool output and organize them as JSON, compute the p95-to-average multiple, and run it three times to fix a baseline that includes the spread.