PCA — Prometheus Certified Associate
The Routing Tree and Burn-Rate Alerts
Goal
You write an Alertmanager configuration yourself from the routing tree to the inhibition rules, create two burn rate alerts derived backward from an SLO, and write a gate that blocks missing runbooks in the build.
Why it matters
The success criterion of an alert is not "did it detect a problem" but "does a person need to do something right now." An alert that does not pass this criterion is harmful even if it is accurate. If 400 accurate alerts ruin the on-call's response speed, accuracy means nothing. So the order is set. Define the SLO from user symptoms, derive the thresholds backward from the budget, combine a long window and a short window with AND to get sensitivity and false-positive control at once, reduce the number of alerts to a level people can read with grouping and inhibition, and enforce a runbook on pages. If you set thresholds in a way like "1.2 times the past maximum," you cannot explain why that value, and a threshold you cannot explain rises a little every time there is an incident.
Steps
- Create
/root/pca-alerting/alertmanager.ymland write the rootroute.receiver: ticket-queue,group_by: ['alertname', 'cluster', 'slo'],group_wait: 30s,group_interval: 5m,repeat_interval: 4h. - In the first item of the root's
routes, put a child route.matchers: ['severity="page"'],receiver: oncall-pager,group_wait: 10s,repeat_interval: 1h. - Define two receivers in
receivers. Give each ofticket-queueandoncall-pagerawebhook_configs, and any internal address starting withhttp://will do forurl. - Write the first item of
inhibit_rules.source_matchers: ['alertname="ClusterDown"', 'severity="page"'],target_matchers: ['severity=~"page|ticket"'],equal: ['cluster']. - Create
/root/pca-alerting/burnrate.ymland write the alertCheckoutErrorBudgetBurnFastunder a group.expris the three termsjob:slo_errors:ratio_rate1h{job="checkout-api"} > (14.4 * 0.001),job:slo_errors:ratio_rate5m{job="checkout-api"} > (14.4 * 0.001), andjob:http_requests:rate5m{job="checkout-api"} > 1joined withand,for: 2m,labelshasseverity: pageandslo: checkout-availability, andannotationshassummaryandrunbook_url. - Add the alert
CheckoutErrorBudgetBurnSlowto the same file.exprisjob:slo_errors:ratio_rate6h{job="checkout-api"} > (6 * 0.001)andjob:slo_errors:ratio_rate30m{job="checkout-api"} > (6 * 0.001)joined withand,for: 15m, andlabelshasseverity: ticketandslo: checkout-availability. - Write
/root/pca-alerting/runbook-gate.sh. It takes the rule file path as its first argument, and if there is even one alert whoselabels.severityispagebut which has noannotations.runbook_url, it prints its name and exits with a non-zero code, and if there is none, it exits with 0.
Notes
- The
exprin steps 5 and 6 may also be written across several lines with anexpr: |block. Grading ignores whitespace. - The verification of step 7 consists of two things: the
burnrate.ymlyou made (which must pass) and a file with a missing runbook that the grader makes (which must be blocked). If you also block alerts withseverity: ticket, it fails. - Hint: extract a list with
yq '.groups[].rules[] | select(...) | .alert' "$1"and, if it is not empty,exit 1. - Common mistake 1: putting
instanceingroup_by. As many alerts arrive as there are Pods. - Common mistake 2: leaving
equalout of the inhibition rule. Alerts from other clusters get inhibited too.
Writing the root route
Create /root/pca-alerting/alertmanager.yml and write the root route. receiver: ticket-queue, group_by: ['alertname', 'cluster', 'slo'], group_wait: 30s, group_interval: 5m, repeat_interval: 4h.
The root's receiver is the default for when no child matches. Remember that what you put in group_by decides the number of alerts. The three timers are, respectively, the wait before the first send, the group update interval, and the repeat interval when nothing changes.
Child routes by severity
In the first item of the root's routes, put a child route. matchers: ['severity="page"'], receiver: oncall-pager, group_wait: 10s, repeat_interval: 1h.
routes is a list under the root and is matched from top to bottom. Unless you turn on continue, it stops at the first match. A child inherits the parent's settings and overrides only the values you state.
Defining receivers
Define two receivers in receivers. Give each of ticket-queue and oncall-pager a webhook_configs, and any internal address starting with http:// will do for url.
If a receiver referenced by name in a route is not in the receivers list, the configuration does not load. This lab environment has no real destination to send to, so you only supply the shape with a webhook URL.
Inhibition rule
Write the first item of inhibit_rules. source_matchers: ['alertname="ClusterDown"', 'severity="page"'], target_matchers: ['severity=~"page|ticket"'], equal: ['cluster'].
Inhibition blocks lower-level symptom alerts while the cause alert is alive. It holds only if the values of the labels written in equal are the same on both sides, so if you leave this label out, alerts from other clusters get inhibited too.
Fast burn alert
Create /root/pca-alerting/burnrate.yml and write the alert CheckoutErrorBudgetBurnFast under a group. expr is the three terms job:slo_errors:ratio_rate1h{job="checkout-api"} > (14.4 * 0.001), job:slo_errors:ratio_rate5m{job="checkout-api"} > (14.4 * 0.001), and job:http_requests:rate5m{job="checkout-api"} > 1 joined with and, for: 2m, labels has severity: page and slo: checkout-availability, and annotations has summary and runbook_url.
The long window decides the burn speed, and the short window decides whether it is still in progress now. Add a minimum traffic gate to this and join the three terms with AND. Since you already use a long window, set for short, just to absorb gaps.
Slow burn alert
Add the alert CheckoutErrorBudgetBurnSlow to the same file. expr is job:slo_errors:ratio_rate6h{job="checkout-api"} > (6 * 0.001) and job:slo_errors:ratio_rate30m{job="checkout-api"} > (6 * 0.001) joined with and, for: 15m, and labels has severity: ticket and slo: checkout-availability.
When the budget is leaking but slowly, it is a matter for a ticket, not for waking someone. Follow the convention of setting the short window to one-twelfth of the long window, and severity must not be page.
Runbook CI gate script
Write /root/pca-alerting/runbook-gate.sh. It takes the rule file path as its first argument, and if there is even one alert whose labels.severity is page but which has no annotations.runbook_url, it prints its name and exits with a non-zero code, and if there is none, it exits with 0.
The script takes the rule file path as its first argument. With yq, pull out the rules whose severity is page and that lack a runbook_url annotation, and if there is even one, print its name and end with failure. You must not also block the ticket class.