TT Lab
Get started
Learn Learning paths Courses

Load Testing

"8% slower" is missing one number: how big the noise is

Continue in TT Lab

In one line

Whether the difference between two runs is a regression cannot be decided by the size of the difference alone. You must first measure the difference that appears when nothing has changed, that is, the size of the noise.

Why this matters

There was a team that attached a performance test to its deployment pipeline. The rule was simple — if p50 gets more than 5% slower than the previous deployment, block the deployment. It looked plausible, and in fact it blocked three times in the first week.

The problem was that two of those three times were false alarms. One change that was rolled back had deleted a single log line, and the other was a comment edit. Later, when the same test was run five times on a commit that changed nothing, p50 spread by as much as 7%. The 5% threshold had been smaller than the noise.

A false alarm is a cost in itself, but the greater damage is that it eats away trust. When a gate misfires often, people learn to turn the gate off or press "retry". And when a real regression goes by, nobody stops it.

Conversely, if you set the threshold too loose, the gate blocks nothing and only gives the illusion that it is blocking. These two failures come from the same root — not having measured the noise when setting the threshold. The cost of measuring the noise is only the time to run the same test a few more times, and that time is always shorter than the time to investigate one wrongly rolled-back change.

How it works

The verdict needs three things. What to measure, how to combine values measured several times into one, and what to compare that one against.

What to measure. The average is pulled around by the tail and the maximum is swayed by a single incident. p50 is stable but cannot see the tail, and p99 sees the tail but swings badly when samples are few. For a gate, use a value with little swing, and it is better to watch the tail separately.

How to combine values measured several times. If you measure once and stop, there is no way to know whether that one was an unlucky run. So you run under the same conditions several times and use the median of those values. The median is not dragged by a single strange run.

What to compare against. This is the heart of it. If you fix the threshold like "5%", nobody knows where that number came from. Instead, you use as the threshold the spread of the values measured repeatedly on the baseline as it is. The simplest form is the maximum minus the minimum, and to make it sturdier you use a multiple of the median absolute deviation. Either way, what matters is that the threshold is a value actually measured on that machine at that time.

This rule has a natural property. If you increase the number of repetitions, the noise estimate improves and you can catch smaller regressions, and if you reduce it to one, the threshold becomes 0 and any difference is reported as a regression. So "how many times to run" is a decision that trades time for sensitivity, not a matter of taste.

What it looks like in the field

The most common failure is measuring the baseline only once. If you take one measured value of yesterday's deployment as the baseline and compare it with one value of today, there is no material at all for deciding the threshold. Then people pull out a fixed percentage, and that number is usually decided in a meeting room.

The second is reading a change in conditions as a regression. If the baseline was measured in the quiet early morning and the new measurement in the daytime with other jobs running alongside, the two numbers are not a comparable pair to begin with. The definition of noise holds only if you measure alternately on the same machine, at the same time of day, with the same load shape.

The third is the baseline going stale. If you measure the baseline once and compare only against that value for months, in the meantime the kernel is upgraded, base libraries change, and machines are replaced. The noise then and the noise now are different values, and the difference from a stale baseline is not a regression but the passage of time. It is safer to re-measure the baseline on the same day, on the same machine, together with the new measurement.

The fourth is judging by eye without a rule. The moment you put two bar charts side by side and say "it's clearly slower", that judgment becomes irreproducible. Even if the next person looks at the same data and reaches a different conclusion, nobody can say they are wrong. If you write the rule down as a script, the verdict remains as a record, and when you later find the threshold was wrong, it is also clear what to fix. And that script must emit an exit code, not a report people read — because only what a pipeline can read can actually block anything.

The fifth is missing an improvement rather than a regression. If you set the gate in only one direction, you cannot know that performance got better either. If a change on the better side greatly exceeds the noise, that must also be recorded — it is usually good news, but sometimes it is a signal that the test has started skipping work.

What you will do in the next lab

You start a target with reproducible jitter added to a base latency and measure five times under the same conditions to get the size of the noise first. Then you measure a version whose latency is 5% longer and a version 50% longer, five times each, set up a rule comparing the medians of the five runs, and judge which exceeds the noise. If you pick the fastest and the slowest of the five baseline runs and imitate a single-run comparison, you see in numbers that even with the same code the gate reports a regression. Finally you build a gate script that takes a baseline directory and a new measurement directory and emits an exit code, and the grader runs that gate with three sets of inputs it made itself to confirm both passes and failures.