TT Lab
Get started
Learn Learning paths Courses

Load Testing

The worst moment disappears from your results

Continue in TT Lab

In one line

A closed-loop load generator does not send requests at all while the server is stalled. Requests that were not sent do not remain in the latency distribution, so the worst moment vanishes from the results entirely. This is called coordinated omission.

Why this matters

I once received a report like this in front of a payment API. "We sent 12,000 requests in one minute and p99 was 60 milliseconds. That is well below the 500-millisecond target." Yet in the server log for the same period, a 2-second stall was recorded once. Both statements came from the same data.

There is nothing strange about it. If the load generator was running at concurrency 1, then while the server held on for 2 seconds, that generator sent nothing because it was waiting for the response. Only the one response that came back after the 2 seconds is recorded as 2 seconds. That is 1 sample out of 200. 1 out of 200 is the 99.5th percentile, so in the p99 slot sits the perfectly fine 50 milliseconds next to it.

Real users do not behave that way. A user does not postpone requests because the server has stalled. They keep coming in during those 2 seconds, and all of those requests pile up in line and wait for nearly 2 seconds. In other words, the report was produced with dozens of requests that actually had a bad experience missing from the measurement.

What makes this flaw scary is that it always fails in the same direction. Coordinated omission does not shake the results at random but makes them wrong only in the direction of looking good. So nobody suspects it, it passes the launch review, and only after an outage do you hear "but it was fine in the test."

How it works

There are two broad ways of applying load.

Model When the next request is sent If the server slows down
Closed loop After the previous request's response The amount sent drops by itself
Open model When a predetermined time arrives The amount sent stays the same, the queue gets longer

A closed loop is the right model for imitating a number of concurrent users. The problem is that this property seeps into the measurement as well. When the server slows down, the load drops with it, so the generator and the server conspire (coordinate) to skip the bad period together. That is where the name comes from.

There are two ways to fix it. The first is to measure with an open model from the start. You fix in advance the intended send time of request i and measure latency as 응답 시각 − 의도한 발사 시각 (that is, the response time minus the intended send time). This value is called corrected latency. If the send was late, the lateness is added to the latency as it is.

The second is to fill in afterward the closed-loop samples you have already measured. This is what HdrHistogram's recordValueWithExpectedInterval does — if one sample's value is larger than the expected interval, it creates and inserts the requests that should have been sent in the meantime but could not be, subtracting the expected interval each time. A single sample with a latency of 2.0 seconds creates 25 more samples at an expected interval of 80 milliseconds: 1.92 seconds, 1.84 seconds, and so on. The reason wrk2 was forked from wrk was also to add this correction.

The results of the two methods point to the same side. So you must do one or the other. However, post-hoc correction has the limitation that you can use it only if you know the expected interval. If the test was run with a target rate set, the expected interval is the inverse of that rate, but in a test run with only concurrency set and no rate, there is no such thing as an expected interval in the first place. There is no way to fix the results of such a test afterward, and you have no choice but to measure again.

What it looks like in the field

The most common picture is a result where "only the maximum is large and p99 is small". When you see this combination, suspect coordinated omission first. A maximum that is thirty times p99 means only one or two samples of the bad period were captured, and that is usually because requests could not be sent in that period.

The second is a result where "I raised the concurrency and p99 actually got better". When concurrency is high, a few requests from other workers do enter and queue even during the stalled period, so a few samples appear, but it is still far fewer than the real arrival rate. The number did not get better; it is just less wrong.

The third is a report that the numbers got worse after changing tools. When you move from a closed-loop tool to an open-model tool, p99 jumps by several times. There have actually been cases where people concluded "the new tool is inaccurate" and reverted. The side that reverted was wrong — the new number was the first correct number.

Finally, after you apply the correction, the SLA verdict often flips. Pass before correction, fail after. What is needed then is not to tweak the threshold but to agree again on which number the promise was made in. If you do not write down in the test report whether correction was applied, half a year later nobody knows which one the number was.

What you will do in the next lab

You start a Python server that holds for 2 seconds only on the 150th request, and measure the same server twice. One is a closed loop with hey, and the other is an open model with a fixed schedule in which 200 requests at 80-millisecond intervals are written out in advance. You recompute p50/p99/max yourself from the two raw files and put them in a table, and implement by hand an HdrHistogram-style correction on the closed-loop samples to build a third distribution. Finally, you set one SLA threshold and confirm the verdict flipping before and after the correction, and leave test design rules in a file so that you never fall into this trap again.