TT Lab
Get started
Learn Learning paths Courses

Observability

The Error Budget Is a Negotiating Tool

Continue in TT Lab

In one line

The period budget is the share used so far, and the burn rate is the speed of recent use. You have to tell the two apart to explain when to allow a deployment and when to stop it.

Why this matters

Suppose you are about to deploy a new feature in a fictional online shop. The last hour has been quiet, but last week's outage has already exceeded the 30-day budget. Conversely, there are days when the monthly budget is ample but payment requests are failing fast right now. If you look at only the recent error rate, you miss the former incident, and if you look at only the monthly number, you cannot respond to the latter. Instead of stopping at a person reading a report, let's build a small decision program that separates the two bases. We assume you know Python dictionaries, conditionals, and JSON input and output.

How it works

First decide whose failures of what kind to count. In this lab's synthetic HTTP data, 5xx counts as failure and the other observed requests count as success. This is the definition for a teaching service, not a claim that every 4xx is always the customer's fault. If a wrong authentication response the service produced or an overload 429 is a user failure, a separate SLI definition is needed. Also document whether to include bot and internal health-check requests, and apply the same scope to the numerator and the denominator.

The allowed error ratio of a request-based 99.9% SLO is 0.001. If total requests over the same 30 days are one million, the allowed error budget is one thousand. If there are 200 failures, the usage ratio is 200/1000=0.2, that is 20%, and at 1200 it is 1.2, exceeding the budget. This calculation needs the failure count and request count for that entire 30 days.

The burn rate is the error ratio of a short window divided by the allowed error ratio. If 200 of 10,000 requests failed in one hour, that is 0.02/0.001=20 times. Even if you compute it over a 12-hour window, it is just the 12-hour burn rate and does not automatically become monthly budget usage. The window length does not change the unit. Even if the recent speed is 20 times, the monthly usage ratio can be 0.2, and even if the current speed is 0, the usage ratio can be 1.2 because of a past outage.

A time-based availability budget is a separate concept. 30 days is 43,200 minutes, so the allowed downtime of a 99.9% time-based SLO is 43.2 minutes. At 99.99% it is 4.32 minutes, that is 4 minutes 19.2 seconds. You must not use this number interchangeably with a request failure count. One minute in a quiet period and one minute when orders are flooding in can have different effects on a request-based SLO. Lab step 2 is a practice in the orders of magnitude of the time budget, and the final program uses the request budget.

Let's also understand the alert threshold of 14.4 in terms of units. Take 30 days as 720 hours and consider a constant rate of using 2% of the total budget in one hour: 0.02×720=14.4. If you assume the remaining budget is intact and the request volume and error ratio are constant, you get an exhaustion forecast of about 50 hours. If there is budget already used, a change in traffic, or old failures that will drop out of the rolling window, this forecast changes. Do not state the future as a fixed number.

Why look at two windows together

Even after an outage ends, the failures remain in the 1-hour average. If the 1-hour burn rate is 20 and the last 5 minutes are 0, you can keep paging by looking only at the long window. A condition that pages only when both 1 hour and 5 minutes exceed 14.4 checks whether the problem is still continuing recently. If you change and to or, this resolve condition disappears. Exactly 14.4 is not an excess.

Be careful with low request counts too. If one of three requests in 5 minutes fails, the ratio is large, but from the number alone you cannot tell whether that one request was an important order or a harmless retry. Do not unconditionally attach a minimum-traffic condition that hides all low-traffic failures. This synthetic experiment does not assume 0 requests is normal but puts it on hold. A real service needs a separate policy that fits its business impact and monitoring for traffic stoppage.

What it looks like in the field

If you fail to fetch the data and convert an empty result to 0, it is mistaken for "no outage". First check that it is the same service and end time and that you collected the entire window you need. The lab Prometheus has about 12 hours of synthetic history. The fact that you wrote a 30-day query against it does not mean 30 days were observed. In the last task you use the separately provided synthetic period aggregate, and do not disguise it as a monthly record extracted from a real Prometheus.

The feature deployment policy of this fictional team has three stages. If the data is unclear, hold; if the period budget usage ratio is 1 or more, freeze; and apart from that, if both the 1-hour and 5-minute burn rates exceed 14.4, freeze. Only the rest is allow. Note the difference that the usage ratio boundary is "or more" and the speed boundary is "exceeds". Both hold and freeze do not proceed with the deployment, but the reasons differ. With a hold, you restore observation first, and with a budget freeze, you prioritize reliability work. Emergency security changes or rollbacks in a real organization need a separate approval policy, and this program does not approve them.

What you will do in the next lab

You query a real Prometheus for the success ratio and the burn rates of the two windows, and check the alert rule. At the end, you write a Python program that reads the synthetic aggregate from stdin. It distinguishes a monthly budget overrun, an ongoing outage, recovery, and missing data, and leaves a report and an exit code. What you need in front of the deploy button at dawn is verifiable evidence rather than a confident-sounding sentence.

Reference: SLO alerting design in SRE, Error budget policy example. The 30 days, decision boundaries, and hold rules above are teaching policies specified in this lab.