PCA — Prometheus Certified Associate
An Alert Succeeds When Action Follows, Not When It Detects
In one line
Alerting rules run the state machine inactive → pending → firing, and Alertmanager processes the alerts it receives in the order Wait → Dedup → Retry → Inhibit → Silence → Notify. If you know where in these two pipelines an alert disappears, you can narrow down "why did the alert not arrive" within 3 minutes.
Why this was needed
The typical path by which on-call collapses goes like this. When an incident happens, the retrospective concludes "we should have known this earlier," and one alert is added. In 2 years there are 300 rules, 400 alerts a day pile up in the Slack channel, and the on-call engineer wakes up three times in the night, does nothing three times, and goes back to sleep.
If 4 of the 400 alerts a day actually need action, the precision is 1%. In this state, the optimal strategy a person learns is "ignore it for now and check later," and because that is rational, it is more dangerous. You cannot reverse it with training or resolve.
So set the practical goal in numbers. At most 2 pages per 12-hour shift, and a page-to-action conversion rate of at least 70%. If you exceed these two, it is time to delete alerts, not add them.
How it works
The alert state machine
inactive --(표현식 매칭)--> pending --(for 경과)--> firing --(매칭 해제)--> resolved
|
+--(중간에 거짓)--> inactive (타이머 0 으로 초기화)
The key to for is consecutive. If it becomes false even once in between, the timer goes back to 0. This is the principle that prevents flapping. There is common bad advice here: "There are many false positives, so let's raise for to 30 minutes." It works, but detection is delayed by 30 minutes, and if an incident ends in 25 minutes, the alert never fires at all. And the real problem remains — the fact that the expression itself is looking at noise.
The right order is to smooth noise with the rate of a long window, apply a short window with AND to confirm it is in progress, and apply for last for only 2–5 minutes to absorb one or two scrape gaps. If you already use a long window and for exceeds 15 minutes, the window design is wrong.
An alert in the firing state is resent to Alertmanager at every evaluation interval. At that time, endsAt is set to the current time + 4 × evaluation_interval, so that it does not expire before the next send arrives. This expiry design is why alerts are automatically resolved when Prometheus dies.
The Alertmanager pipeline
| Stage | What it does | Easy to miss |
|---|---|---|
| Wait | When a group is newly created, collects for group_wait |
Default 30 seconds |
| Dedup | Checks with the notification log whether the alert was already sent | In HA, this log is shared by gossip |
| Retry | Exponential backoff retry on failure | |
| Inhibit | If the source is firing, blocks the target | Holds only if the equal labels are the same |
| Silence | Matches against silences people created | A silence with no expiry is a sign of an alert that should be deleted |
| Notify | The actual send |
The difference among the three timers is an exam regular. group_wait is the wait before the first send of a new group (default 30 seconds), group_interval is the send interval when a new alert is added to an existing group (default 5 minutes), and repeat_interval is the interval for reminding about a group with no change (default 4 hours).
What you put in group_by decides the number of alerts. If you group by instance, as many alerts arrive as there are Pods, and you have to group by service or SLO to get a number a person can read.
Inhibition and silences are often confused. Inhibition is rule-based automatic blocking written in the configuration file, and a silence is a temporary block people create through the UI or API. Both block alerts, but they are still visible in the UI.
The multi-window burn rate is the standard combination in the SRE workbook.
| Budget consumed | Long window | Short window | Threshold burn rate | Response |
|---|---|---|---|---|
| 2% | 1 hour | 5 minutes | 14.4 | Page immediately |
| 5% | 6 hours | 30 minutes | 6 | Page immediately |
| 10% | 1 day | 2 hours | 3 | Ticket |
| 10% | 3 days | 6 hours | 1 | Ticket |
By convention, the short window is set to one-twelfth of the long window. The role of the short window is not sensitivity but recovery speed. Without a short window, even after the incident ends, the long window keeps the alert going for the length of the window.
What it looks like in the field
A single rule that left out for: once caused more than 50 alerts a day to pour in. The result was as expected — people started ignoring those alerts, and once ignoring became a habit, the real signals in the same channel were buried along with them. It was not the accuracy of one alert but the response speed of the whole on-call that was ruined.
Low-traffic periods are also a recurring trap. At dawn, when 3 requests come in over 5 minutes and 1 fails, the error rate becomes 33% and crosses every burn rate threshold at once. Phantom alerts disappear only if you apply a minimum traffic condition with AND. If traffic is 0 altogether, the ratio becomes NaN and the alert quietly vanishes, and in that case a separate throughput-drop alert has to catch it.
And enforce a runbook on pages. A person woken at 3 a.m. is not in a state to exercise creativity. What is needed is a document that lists three things to check and two actions to take. If you use a CI gate to find rules that have severity=page but no runbook_url and break the build, this discipline no longer depends on people's memory.
What you will do in the next lab
Under /root/pca-alerting/, you write an Alertmanager routing tree and inhibition rules, and create two burn rate alerts, a fast burn and a slow burn. At the end, you write a CI gate script that catches missing runbooks and check that it passes a good file and blocks a file without a runbook.