Enough Accurate Alerts Add Up to an Inaccurate One
In one line
The quality of an alert is measured not by accuracy but by actionability. An alert that leaves nothing for a person to do each time it fires becomes a lie as a whole, even if each one is accurate.
Why this matters
Adding alerts is easy. After every incident, a metric appears that makes you think "if only we had known this earlier", and if you put a threshold on that metric, you have something to say at the next meeting. After a year of that, 80 alerts pile up in the rules file.
The problem comes next. The on-call engineer's phone goes off twenty times in the night, and nineteen of them turn out in the morning to be nothing. People learn — from then on, they do not get up right away even when they see an alert. And then the twenty-first, a real outage, arrives with the same sound.
This is alert fatigue. No matter how much you raise the accuracy of individual alerts, it does not get solved. The problem is the total volume of pages. The Google SRE Book puts it this way: "the output of a monitoring system must be manageable by humans". Manageability is decided not by each rule but by the sum.
How it works
The skeleton of the solution is two things.
First, alert on symptoms and investigate by causes. What users experience — requests fail, are slow, or give wrong results — comes down to only a few things. The causes that produce those symptoms, on the other hand, number in the dozens. If you attach an alert to every cause you get dozens of pages, but if you attach alerts to symptoms it shrinks to a few. Cause metrics are not removed but moved down to dashboards and investigation tools. What you look at after a person has woken up and what wakes a person up are different places.
Second, give alerts a duration. If an alert fires as soon as a condition becomes true once, a momentary spike of a value becomes a page as it is. If you set for to 5 minutes, only what has held for more than 5 minutes fires. With this one knob, it is common for the number of pages to drop to a single digit — however, detection is delayed by that much, so how much delay is acceptable is answered by the error budget.
The question to ask when choosing alerts can be reduced to one. Is there something the person who received this notification can do right now? If not, it is not a page but a ticket or a dashboard panel.
| Nature | Where to send | Example |
|---|---|---|
| A person must step in now | Page | The error ratio of user requests is exceeding the objective |
| It can be handled within this week | Ticket | The disk will fill in 6 hours |
| Look at it when investigating | Dashboard | p99 latency, queue depth, per-instance CPU |
There are places where duration alone is not enough. If you wait for a condition to hold for 5 minutes, you miss small incidents and learn about big ones late. That is why SLO-based alerts use burn rate instead of duration — you look together, with a short window and a long window, at how fast the budget is being consumed at the current speed, and fire only when both exceed at the same time. The structure is that the short window handles fast detection and the long window handles noise removal.
It is good to set a budget for alerts. It is a promise such as that one person will not receive more than two pages per week. With a budget, each time you add a new alert you also ask "what will we remove?", and that question slows the speed at which the rules file grows. When the budget is exceeded, the answer is often to fix the system rather than fix the rules — an alert that fires often is usually honestly reporting a failure that happens often.
And two things are enforced for every alert. One is a runbook link. A person who receives an alert they see for the first time at three in the morning must be able to learn from that one link what to check and what to roll back. The other is an owner. An alert that nobody is responsible for stays forever because nobody can delete it. A team that holds an alert review every quarter, pulls out a list of "alerts that never fired in the last 90 days, or fired but led to no action", and deletes them keeps its rules file small.
What it looks like in the field
One team's rules file had an alert that said "the request count is lower than usual". The intention was good — if traffic stops, it could be an outage. But at dawn traffic is naturally low. That alert was true for 609 minutes out of 12 hours, that is, nearly half, and the on-call engineers had set up a filter to ignore that notification automatically. It was not merely useless; it was harmful in that it cultivated the habit of creating filters.
There is a case on the other side too. An alert that fires when p99 latency exceeds 0.5 seconds had no duration, so it fired sixty-five times in 12 hours. When for: 5m was added it became eight times, and with for: 15m it became one. Not a single character of the rule's condition changed. All that changed was the team's agreement on "how long must it persist to be worth waking someone".
What you will do in the next lab
You write four cause-based alerts as a rules file and build a tool that counts how many times each would have fired against the backfilled 12 hours of data. You change the duration to 0, 5, and 15 minutes to measure how the number of pages decreases, and compare it with one symptom-based alert. Then you count how many minutes, of the time a cause alert fired, users experienced nothing, decide in a table which to keep as pages and which to move down based on those numbers, and finally create a rules file with only the paging rules and get it through the check.