TT Lab
Get started
Learn Learning paths Courses

Observability

Change the Denominator and Availability Changes

Continue in TT Lab

In one line

An SLI is not a query but a definition. Deciding what counts as a good event and what counts as a valid event determines most of the number.

Why this matters

The sentence heard most often in an outage meeting is "what percent was our availability?" Yet given the same 12 hours of data, three people bring 98.9%, 97.7%, and 97.2%. All three calculations are correct. They just answered different questions.

The first person counted only 5xx responses as failures. The second counted 400-range responses as failures too. The third counted not requests but minutes — if the error ratio exceeded a threshold during a minute, that whole minute was treated as a bad minute.

This difference is not wordplay; it costs money and people's time. If the definition is loose, it looks as if error budget remains and a risky deployment gets through, and if the definition is too tight, people are woken up on nights when nothing happened. So before deciding what percent to set the objective (SLO) at, you must first agree on what to count.

How it works

The Google SRE Workbook writes an SLI in one sentence — the ratio of good events to valid events. So when you write a definition, you must fill in three fields.

Field What it decides Common mistake
Event What to count as one Mixing requests and time windows
Good event What to treat as success Looking only at the status code and counting a slow success as a success
Valid event What goes into the denominator Putting health checks, bots, and internal calls in the denominator

Changing even one of the three fields changes the number. The denominator is especially scary. If you put a health check that runs six times per second into the denominator, the failures users experience are diluted by health check successes and the outage is pushed below the decimal point. Conversely, if you narrow the denominator to a single user journey, the same incident looks much bigger.

Changing the unit of an event from requests to time changes the nature itself. A request-based measure gives a bigger penalty to incidents in high-traffic hours. A time-based measure counts a 20-minute outage at 2 a.m. and a 20-minute outage at 2 p.m. equally as 20 minutes. Which is right depends on what the service promised its users.

When you put latency into the success condition, you must know the limits of histograms. The request count counter has a status code label, but the latency histogram usually does not. Then you cannot count "requests that returned 200 and finished within 100 milliseconds" in one query. To multiply the two signals, you have to fix the instrumentation so that the metrics are exported that way — this is the moment when the definition decides the instrumentation.

The place of measurement is also part of the definition. The number differs depending on whether you measure the same request inside the application, at the front proxy, or in the browser. If you measure inside the application, cases where the connection dropped and the response never reached the user remain as successes, and if you measure at the proxy, the time the proxy itself was down disappears from the record entirely. Wherever you measure, you must also write down which failures that place cannot see, so that you can trust the number later.

You also have to decide how to cut the period. A calendar basis (the budget resets on the 1st of each month) is good for contracts but leads to risky deployments being crammed in at the end of the month. A rolling window (the last 30 days) prevents that kind of last-minute cramming, but past incidents follow you around for a month. There is no choice that is right in both cases; you choose by which one changes the team's behavior in the direction you want.

Finally, do not create several SLIs and average them into one. Availability, latency, and correctness are different promises, so averaging them creates a state in which none of the promises is kept while the number alone looks good. It is much easier to handle if you measure each separately, set an objective on each, and burn down each budget separately. The demand to bundle several SLIs into one number usually means wanting to make the report shorter, and the moment you grant that demand, nobody can tell which promise was broken.

What it looks like in the field

Once, this happened at a payment service. The dashboard availability was above 99.9% but customer inquiries kept coming in. Internal batch calls were in the denominator, and those calls were half of the total. When the denominator was narrowed to user requests only, the availability for the same period dropped to 99.2%. The number did not get worse; it was only then that the correct number came out.

Accidents in the opposite direction are common too. One team set the principle that "slow is also failure" and set the threshold at 100 milliseconds. At that moment availability sank to 85%, the error budget was used up every morning, and deployments were blocked indefinitely. This was because the threshold came not from what users feel but from the owner's preference. When you change a definition, you must first re-measure the past 30 days with that definition and confirm that you can live with that number.

What you will do in the next lab

You calculate four SLIs yourself from the 12 hours of data in the Pod's Prometheus. On a request basis, you count only 5xx as failures, then count 400-range as failures too, then put a 100-millisecond latency into the success condition, and finally count 1-minute windows as events. You see how far the four numbers diverge and convert them into the 30-day allowed downtime, then pick one and declare it in a machine-readable format, and decide whether to allow or block a deployment under two objectives.