Some Alerts Are Accurate and Still Harmful
In one line
The success criterion of an alert is not "did it detect a problem" but "does a person need to do something right now". An alert that fails this criterion is harmful even if it is accurate.
Why this matters
If you add one rule after every incident, after two years you have 300 rules and 400 messages a day piling up in Slack. The on-call person wakes up three times in the night and three times does nothing and goes back to sleep. If 4 of 400 alerts a day led to real action, the precision is 1%.
In this state, the optimal strategy people learn is "ignore it for now and check later". It is more dangerous precisely because it is rational. It is not a problem that adding one more alert solves.
How it works
There is one principle. Page only on symptoms that users experience. CPU at 90% is not a symptom. If CPU is 90% and response time is normal, nothing has happened, and if CPU is 40% and responses take 5 seconds each, it is a serious incident. If you page on a cause metric, both cases come out wrong.
What to page on: availability, latency (the ratio exceeding a threshold), a sharp drop in throughput, freshness, correctness. What not to page on: CPU, memory, disk IOPS, Pod restart counts, thread pool utilization, GC time, replica count. There are only two exceptions. Things that cannot be undone and need lead time (disk headroom, certificate expiry, quota exhaustion), and failure of the observability system itself (up == 0 or absent(up)). Without this alert, the other alerts die silently.
You must also be able to explain a threshold. Where did the 5% in "alert if the error rate exceeds 5%" come from? Mostly it came from nowhere. A threshold you cannot explain creeps up a little after every incident.
for: fires only if the expression stays true continuously for that long, and if it is false even once in between, the timer resets to 0. So "there are many false positives, so let's raise for to 30 minutes" is an anti-pattern. Detection is delayed by 30 minutes, and if the incident ends in 25 minutes, the alert never fires at all. The right order is ① smooth with a long-window rate ② confirm it is still ongoing with a short-window AND ③ for is only 2–5 minutes, at the very end.
What it looks like in the field
Routing has three levels. Page (must wake up now), ticket (business hours), dashboard (creates no alert). There is no fourth level. The compromise of "let's just send it to the Slack channel for now" is the worst — nobody is responsible and nobody turns it off.
The most common mistake in grouping is putting instance in group_by. You get as many alerts as there are Pods. And if severity=page has no runbook_url, CI should break the build. A person woken up at 3 a.m. is not in a state to exercise creativity. However, a runbook that has only a link and empty contents is worse than having no link.
Set targets as numbers. On a 12-hour shift, an average of 2 pages or fewer, and a rate of pages leading to action of 70% or more. Once you go beyond these two numbers, it is time to delete alerts, not add them. Downgrade or delete pages that have not led to action for 3 months in a row, and delete alerts whose silence has lasted 30 days or more. A silence is the most honest signal that an alert is wrong.
What an alert should contain
What a person woken at 3 a.m. reads on the screen is the one line of the alert title. That one line and the body must produce the next action, so alerts must contain set items.
- What is wrong, in the user's words. "Payment API error rate 12%", not "hikari_pending > 0".
- How bad it is. Write the current value together with the threshold. The response differs depending on whether it is 12% or 51%.
- Since when. Something that just started and something at its 40th minute are different situations.
- Where. The service, the environment, and as much scope as needed. However, if you put in even the Pod name, the alert splits by the number of Pods, as mentioned earlier.
- What to do next. A runbook link and a link that goes straight to the dashboard.
One more thing is recommended. A way to record "this alert fired but no action was needed". If the on-call person can leave that with a single click on the spot, then weeks later, which alerts are noise comes out as data. You reach a conclusion far faster than by arguing from memory in a retrospective.
What you must decide together whenever you create an alert is the condition for deleting it. If, when adding a rule, you also write "delete it if it does not lead to action within 3 months", the situation where there are 300 rules never arises in the first place. Rules are easy to add and hard to delete, so the grounds for deleting must be created at the same time.
Finally, you must test the alert itself. After writing a rule, artificially create its condition and confirm that it actually fires, and that it resolves when the situation ends. An alert that does not resolve is as bad as an alert that does not fire. If it fires once and then stays red forever, nobody notices when the same thing really happens next time.
What you will do in the next lab
You classify symptoms and causes, fill in the required fields of alert rules, define routing, inhibition, and grouping, fill in a runbook for real, and finally build a checker yourself that catches defective rules in CI.