SLOs — Deciding How Much Breakage Is Allowed
Compute the Budget and Test the Alert
Goal
"99.9% availability" is easy to write into a contract, but few people know how many minutes per month that is. And most alerts have nothing to do with the SLO.
In this lab you work out the numbers yourself, rewrite the alert with those numbers, and test that the alert really fires.
Getting started
cp -r /opt/lab/slo/* .
python3 budget.py 99.9
promtool check rules rules.yml
promtool test rules test.yml
Files
| File | What it does |
|---|---|
budget.py |
Computes error budgets and burn rates. Do not edit it |
rules.yml |
The alerting rules. Edit this one |
test.yml |
Unit tests for the rules. Edit this one too |
Time series notation
0+90x20 in test.yml means start at 0 and increase by 90 every minute, 20 times.
That is, 90 requests per minute.
Steps
- Error budget →
01-budget.txt - Burn rate →
02-burn.md - What is wrong with today's alert →
03-why-bad.md - Rewrite with burn rate →
rules.yml - Test that it fires →
test.yml - Test that it stays quiet →
test.yml - Two windows →
07-window.md - Wrap-up →
08-notes.md
Note
For steps 4, 5, and 6, the grader runs promtool directly to check.
How many minutes a month is 99.9%
Work out the monthly error budget for each of 99.9, 99.95, and 99.99, and save it to 01-budget.txt.
Write it like python3 budget.py 99.9.
The goal is not to memorize the numbers but to build a feel for the orders of magnitude. 99.9% and 99.99% differ by one digit in how they are written, but the allowed time differs by a factor of ten.
Before you write 99.99% into a contract, you need to know how many minutes per month that is.
When will it run out at the current speed
For each of burn rates 1, 6, and 14.4, work out when the budget runs out and write it in 02-burn.md, and explain where the number 14.4 comes from.
python3 budget.py 99.9 --burn 14.4.
Burn rate is how many times the allowed error rate the error rate is. At 1x you use up exactly the whole budget in a month (as designed).
14.4 comes from this — it is the speed at which you burn 2% of a month's budget in 1 hour. 0.02 × 30일 × 24시간 = 14.4. At this speed, you need to wake someone up now.
Why today's alert is useless
Read the current rule in rules.yml and write in 03-why-bad.md why it has nothing to do with the SLO. Give one example where it fires when it should be quiet and one where it stays quiet when it should fire.
The current rule is "5-minute error rate above 1%." It has nothing to do with the SLO (0.1% allowed).
Think about this — what if an error rate of 0.5% continues all month? (That is 5 times the budget, yet it stays quiet.) Conversely, what if 1.2% spikes for 5 minutes in the early morning? (The budget barely moves, but it wakes someone up.)
Alert fatigue comes not from having many alerts but from having many useless alerts.
Rewrite it with burn rate
Edit rules.yml to turn it into an alert based on burn rate. promtool check rules rules.yml must pass.
The threshold is 14.4 × (1 - SLO). If the SLO is 99.9%, it is 14.4 * 0.001.
expr: |
(sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))) > (14.4 * 0.001)
Change the alert name too — it is no longer "the error rate is high" but "the budget is being burned quickly." The name decides the response.
Test that the alert really fires
Edit test.yml so that promtool test rules test.yml passes. It must match the new alert name and the comments.
This is the most valuable step in the lab. Few people attach unit tests to alerting rules. That is why alerts that do not fire when the outage actually comes are common.
0+90x20 in input_series means increase by 90 every minute, 20 times. If you set errors to 0+10x20, that is a 10% error rate, so it fires under any criterion.
The grader runs promtool test rules directly.
Also test that it stays quiet when it should not fire
Add one more case to test.yml where the alert must not fire. It is a normal state running within the budget.
Express "there must be no alerts at all" with exp_alerts: [].
An error rate of about 0.05% (for example 0+1x20 against 0+2000x20) is within the SLO.
Testing only that it fires when it should is half the job. You prevent alert fatigue only by also testing that it stays quiet when it should.
The fast window and the slow window
Write in 07-window.md why a single window is not enough, and propose a setup that uses two windows.
With only a short window (5 minutes), you get woken up by brief spikes. With only a long window (1 hour), you learn about fast burning late.
So you usually join the two with and — the alert fires only when the short window and the long window both exceed the threshold. Then momentary spikes are filtered out and real burning is caught quickly.
And you split by severity — fast burning (14.4x) gets an immediate page, slow burning (6x or lower) gets a ticket.
Wrap up
Write at least three lines in 08-notes.md: what burn rate is, why you should test alerting rules, and what the team decided to do when the budget runs out.
The text must include 번 레이트, 시험, and 예산 (the Korean words for burn rate, test, and budget). The last item is the hard one — if no response is defined for budget exhaustion, the SLO is decoration.