Pushing Alerts Through Alertmanager
Goal
You send alerts through yourself to see how alerts generated by Prometheus are split, grouped, and inhibited as they pass through Alertmanager.
Why it matters
Alert design fails in two directions. Either it fires so much that nobody looks, or it does not fire and an outage is missed. Both give the same result — the alerts lose trust.
That is why routing, grouping, and inhibition are needed. If one cause produces twenty alerts, people do not read all twenty. One inhibition rule reduces them to one.
Steps
- Start Alertmanager
- Split routes by severity →
/etc/alertmanager.yml - Define two receivers
- A rule where critical inhibits warning
- Check with
amtool check-configand apply - Push an alert in directly with
amtool alert add - Try creating a silence
- Attach a runbook link to the alert
Notes
amtoolhas/root/.config/amtool/config.ymlpreinstalled, so you can simply useamtool alertwithout--alertmanager.url.amtool check-configchecks only the syntax. To see whether routing goes as intended, useamtool config routes test severity=critical.- If you changed the configuration, you must restart Alertmanager for it to take effect.
Start Alertmanager
Start Alertmanager
Start it with lab-start-alertmanager. 9093 must show as [O] in lab-status. The alertmanagers target is already set in the Prometheus configuration.
Write the route tree
Split routes by severity → /etc/alertmanager.yml
Edit /etc/alertmanager.yml. Under route, use routes to split severity=critical and severity=warning and send them to different receivers. Give critical a short group_wait and warning a long one — so that the urgent one is announced first.
Define two receivers
Define two receivers
Put at least two in the receivers list. You do not need a real destination — a receiver with only a name is also a valid configuration. What you learn here is not 'where to send' but 'what to split and send'.
Inhibition rule
A rule where critical inhibits warning
Add inhibit_rules. While a critical fires for the same target, the warning is suppressed. You need three things, source_matchers / target_matchers / equal, and equal is the definition of 'the same target' — if you leave it out, even unrelated alerts get buried.
Configuration check
Check with amtool check-config and apply
Check with amtool check-config /etc/alertmanager.yml. If it passes, restart Alertmanager (pkill alertmanager; lab-start-alertmanager) to apply it.
Push in an alert directly
Push an alert in directly with amtool alert add
Create one alert with amtool: amtool alert add TestAlert severity=critical service=shop-api. Then check that it arrived with amtool alert.
Create a silence
Try creating a silence
Bury it for a while with amtool silence add alertname=TestAlert -d 1h -c '점검 중'. Check it with amtool silence query. If you use a silence instead of turning off alerts during maintenance, it comes back to life automatically afterward — there is no accident of turning it off and forgetting.
Attach a runbook link to the alert
Attach a runbook link to the alert
Add annotations.runbook_url to the alert in /etc/prometheus/rules/slo.yml. What a person woken at 3 a.m. needs is not 'what is wrong' but 'what to do'. Check with promtool and then reload.