TT Lab
Get started
Learn Learning paths Courses

Observability

Pushing Alerts Through Alertmanager

Continue in TT Lab

Goal

You send alerts through yourself to see how alerts generated by Prometheus are split, grouped, and inhibited as they pass through Alertmanager.

Why it matters

Alert design fails in two directions. Either it fires so much that nobody looks, or it does not fire and an outage is missed. Both give the same result — the alerts lose trust.

That is why routing, grouping, and inhibition are needed. If one cause produces twenty alerts, people do not read all twenty. One inhibition rule reduces them to one.

Steps

  1. Start Alertmanager
  2. Split routes by severity → /etc/alertmanager.yml
  3. Define two receivers
  4. A rule where critical inhibits warning
  5. Check with amtool check-config and apply
  6. Push an alert in directly with amtool alert add
  7. Try creating a silence
  8. Attach a runbook link to the alert

Notes

Start Alertmanager

Start Alertmanager

Start it with lab-start-alertmanager. 9093 must show as [O] in lab-status. The alertmanagers target is already set in the Prometheus configuration.

Write the route tree

Split routes by severity → /etc/alertmanager.yml

Edit /etc/alertmanager.yml. Under route, use routes to split severity=critical and severity=warning and send them to different receivers. Give critical a short group_wait and warning a long one — so that the urgent one is announced first.

Define two receivers

Define two receivers

Put at least two in the receivers list. You do not need a real destination — a receiver with only a name is also a valid configuration. What you learn here is not 'where to send' but 'what to split and send'.

Inhibition rule

A rule where critical inhibits warning

Add inhibit_rules. While a critical fires for the same target, the warning is suppressed. You need three things, source_matchers / target_matchers / equal, and equal is the definition of 'the same target' — if you leave it out, even unrelated alerts get buried.

Configuration check

Check with amtool check-config and apply

Check with amtool check-config /etc/alertmanager.yml. If it passes, restart Alertmanager (pkill alertmanager; lab-start-alertmanager) to apply it.

Push in an alert directly

Push an alert in directly with amtool alert add

Create one alert with amtool: amtool alert add TestAlert severity=critical service=shop-api. Then check that it arrived with amtool alert.

Create a silence

Try creating a silence

Bury it for a while with amtool silence add alertname=TestAlert -d 1h -c '점검 중'. Check it with amtool silence query. If you use a silence instead of turning off alerts during maintenance, it comes back to life automatically afterward — there is no accident of turning it off and forgetting.

Attach a runbook link to the alert

Attach a runbook link to the alert

Add annotations.runbook_url to the alert in /etc/prometheus/rules/slo.yml. What a person woken at 3 a.m. needs is not 'what is wrong' but 'what to do'. Check with promtool and then reload.