SLOs — Deciding How Much Breakage Is Allowed
Alerting on Burn Rate
One-line summary
If you alert on burn rate, the alert fires when it should and stays quiet when it should.
99.9% is 43 minutes a month
| SLO | Allowed per month (30 days) |
|---|---|
| 99% | 7.2 hours |
| 99.9% | 43.2 minutes |
| 99.95% | 21.6 minutes |
| 99.99% | 4.32 minutes |
99.9 and 99.99 differ by a single digit in how they are written, but they are a factor of ten apart. Look at this table before you write 99.99% into a contract — with 4 minutes a month, a single bad deployment can use it all up.
An error budget is meant to be spent
If your target is 99.9%, the 0.1% is yours to spend. Leaving it unspent is nothing to be proud of — if budget keeps being left over, either the target is too low or you are deploying too rarely.
This view is the heart of an SLO. Instead of aiming for perfection, you agree on how much breakage is acceptable.
Burn rate
How many times the allowed error rate the current error rate is is called the burn rate.
- 1x → you use up exactly the whole budget by the end of the period (as designed)
- 14.4x → you use it all up in 2.08 days
The number 14.4 comes from this — it is the speed at which you burn 2% of a month's budget in 1 hour.
0.02 × 30일 × 24시간 = 14.4
At this speed, you need to wake someone up now. At around 2x, on the other hand, a ticket is enough.
Why today's alert is useless
Take one common rule.
5분 오류율 > 1%
The SLO is 99.9% (0.1% allowed), but this rule did not come from that. So both of these happen.
- It fires when it should be quiet — 1.2% spiked for 5 minutes in the early morning. The budget barely moved, but a person was woken up
- It stays quiet when it should fire — 0.5% persists all month long. That is 5 times the budget, yet it is below the threshold, so nothing happened
Alert fatigue comes not from having many alerts but from having many useless alerts.
Rewritten with burn rate
오류율 > 14.4 × (1 - SLO)
If the SLO is 99.9%, the threshold is 0.0144 (1.44%). The numbers alone look similar, but the meaning is different — this says "2% of the budget is being burned per hour," and so what to do follows from it.
You also change the alert name. Not HighErrorRate but ErrorBudgetBurnFast. The name decides the response.
One window is not enough
- Short window only (5 minutes) → you get woken up by brief spikes
- Long window only (1 hour) → you learn about fast burning late
So you join the two with and. The alert fires only when the short window and the long window both exceed the threshold. Momentary spikes are filtered out, and real burning is caught quickly.
You also split by severity.
| Burn rate | Windows | Response |
|---|---|---|
| 14.4 | 5 minutes + 1 hour | Page immediately |
| 6 | 30 minutes + 6 hours | Page immediately |
| 3 | 2 hours + 1 day | Ticket |
| 1 | 6 hours + 3 days | Ticket |
What should you measure the SLI with
The burn-rate calculation is arithmetic, but unless you first settle what counts as success, the number means nothing. Three things are decided here.
Where to measure. If you measure in server logs, you cannot see requests cut off at the load balancer; if you measure at the load balancer, you cannot see clients' DNS and TLS failures. The closer you are to what users experience, the better, but the more failures you do not control get mixed in. A practical combination is to use the load balancer as the default and synthetic monitoring running from outside as a supplement.
What to count as a failure. 5xx is clear, but 4xx is ambiguous. 400 is usually the client's fault, so leave it out, and it is more honest to include 429, since we were the ones who blocked it. If you exclude 404, you miss an outage where routing breaks and everything returns 404. Write the definition down, and when you change it, recompute past figures as well.
Slow is also a failure. A response that succeeded after 30 seconds is a failure to the user. So a latency SLI is set not as an average but as the proportion of requests that came in within a threshold. If you write "the proportion of requests that came in within 300 ms is 99%" rather than "95% of requests are within 300 ms," you can compute the budget the same way as for availability.
sum(rate(http_request_duration_seconds_bucket{le="0.3",job="api"}[5m]))
/
sum(rate(http_request_duration_seconds_count{job="api"}[5m]))
Do not count every request with the same weight. If health checks and bot traffic are half the denominator, the failures users experience barely show up in the number. Conversely, if a single batch API sends thousands of requests per second, that one client dominates the SLI. Measuring per user journey is the answer, but it costs effort, so at the very least measure one critical path separately.
An SLO is a decision tool, not a contract. You first need an agreement that when budget remains you may make risky changes, and when it is used up you focus on stabilization. Without that agreement, measuring the number only gives you material to argue over after an outage.
Test your alerts too
This is the part most often skipped. promtool supports unit tests for alerting rules.
promtool test rules test.yml
You feed in fake time series and assert "this alert must fire at this point." Without this, you will deploy alerts that do not fire when the outage actually comes — and you will find out only during the outage.
You also test that it stays quiet when it should not fire (exp_alerts: []). Testing only when it fires is half the job.
When the budget runs out
If you have no answer here, the SLO is decoration.
Agree on these in advance —
- Stop deploying new features and spend the time on stabilization
- Postpone risky changes until the budget recovers
- Raise the recurring cause to the top priority of the next sprint
This agreement is harder than setting the number, and without it the number does nothing.
In the field
The first thing you run into when introducing this approach is not technology but agreement. Writing down 99.9% is easy, but deciding in advance what to stop when half the budget is left means talking with the product side, and that conversation is not always comfortable.
That is why it is better to start with a period of observation only, without setting alerts. After you look at a month or two of real burn curves, it becomes clear whether 99.9% is the right number for our service, or whether nobody complains even at 99.5%. If you set a number you cannot keep, the budget runs out in the first week of every month, and then nobody looks at that number anymore.
And what to measure the SLI with is a bigger decision than it seems. If you measure at the load balancer, an application crash shows up as 5xx but client-side failures are invisible; if you measure in the application, it is the other way around. Writing one line in the docs about where you are measuring greatly reduces arguments over the numbers later.