TT Lab
Get started
Learn Learning paths Courses

Observability

Seventy Alerts Overnight and Users Noticed Nothing

Continue in TT Lab

Goal

You count four cause-based alerts and one symptom-based alert on real data, and decide what to send to people and what to move down, based on the number of pages and the time of pointless pages.

Why it matters

Adding alerts is easy and reducing them is hard. The reason it is hard is that there is no evidence — "isn't this alert noisy?" is an opinion, but "this alert fired sixty-five times in 12 hours and 230 minutes of that was time when users experienced nothing" is data. With data, the meeting gets shorter. And the same data also decides the value of the duration knob. Alert design is not about choosing thresholds but about choosing what to detect within a total volume of pages that people can manage.

Steps

  1. Write a Prometheus rules file to /root/obs-alert-symptom/rules/causes.yml. The group name is causes and it contains four alerts. CauseQueueDepth fires when queue_depth{job="shop-api",queue="orders"} exceeds 100, CauseDiskPredict fires when the 6-hour-ahead forecast from the last 1 hour's trend of node_filesystem_avail_bytes{job="node"} is below 0, CauseLatencyTail fires when p99 latency over the last 5 minutes exceeds 0.5 seconds, and CauseTrafficLow fires when the requests per second over the last 5 minutes is below 55. It must pass promtool check rules. Do not add for: in this step.
  2. Create /root/obs-alert-symptom/count.py. When called as python3 count.py '<경보 식>' <for 분> (the alert expression and the for duration in minutes), it scans the last 12 hours at 1-minute intervals and prints one line, episodes=<호출 수> minutes=<울린 분> (the number of pages and the minutes fired). When the minutes in which the condition is true continue for L minutes in a row, if L is for + 1 or more, count it as one page, and the time fired is L - for minutes. If for is not given, treat it as 0.
  3. Create /root/obs-alert-symptom/counts.tsv. It has four lines with no header, and each line has three tab-separated columns, <경보이름> <호출 수> <울린 분> (the alert name, the number of pages, and the minutes fired). Leave for at 0, and write the four alerts you wrote in the step 1 rules file by their names as they are. Get the values with the tool from step 2.
  4. Pick the alert that fired the most in step 3, and count again while changing for to 0, 5, and 15 minutes. In /root/obs-alert-symptom/for-effect.tsv, write three lines with no header, and each line is <for 분> <호출 수> <울린 분> (for minutes, number of pages, minutes fired).
  5. In /root/obs-alert-symptom/rules/symptom.yml, write a symptom group and one alert, SymptomErrorRatio. The condition is that the 5xx response ratio over the last 5 minutes exceeds 1%, and attach for: 5m and a runbook_url annotation. After it passes promtool check rules, count the same expression with for at 0 and write one line, episodes=<수> minutes=<분> (number and minutes), to /root/obs-alert-symptom/symptom.txt.
  6. Create /root/obs-alert-symptom/overlap.py. When called as python3 overlap.py '<원인 식>' '<증상 식>' (the cause expression and the symptom expression), it scans the last 12 hours at 1-minute intervals and prints one line, cause_minutes=<원인이 참인 분> outside_minutes=<그중 증상이 참이 아닌 분> (the minutes the cause is true, and of those the minutes the symptom is not true). Measure each of the four cause alerts with that tool and write four tab-separated three-column lines, <경보이름> <원인 분> <증상 밖 분> (alert name, cause minutes, minutes outside the symptom), to /root/obs-alert-symptom/falsepages.tsv.
  7. Write five lines in /root/obs-alert-symptom/triage.tsv. Each line has three tab-separated columns, <경보이름> <page|ticket|dashboard> <근거> (alert name, the classification, and the grounds), and you write all four cause alerts and SymptomErrorRatio. There are two rules — the symptom alert must be page, and a cause alert whose minutes outside the symptom exceeded 60 in step 6 cannot be set as page. The grounds must include at least one number obtained in the previous steps and be at least 20 characters.
  8. In /root/obs-alert-symptom/rules/page.yml, include only the alerts you classified as page in step 7. The group name is page, and each alert must have for: of 5 minutes or more, severity: page in labels, and runbook_url in annotations. After it passes promtool check rules, count those alerts with each one's own for value and write the sum of the page counts as one line, pages_after=<수> (the number), to /root/obs-alert-symptom/after.txt.

Notes

Write four cause-based alerts as a rules file

Write a Prometheus rules file to /root/obs-alert-symptom/rules/causes.yml. The group name is causes and it contains four alerts. CauseQueueDepth fires when queue_depth{job="shop-api",queue="orders"} exceeds 100, CauseDiskPredict fires when the 6-hour-ahead forecast from the last 1 hour's trend of node_filesystem_avail_bytes{job="node"} is below 0, CauseLatencyTail fires when p99 latency over the last 5 minutes exceeds 0.5 seconds, and CauseTrafficLow fires when the requests per second over the last 5 minutes is below 55. It must pass promtool check rules. Do not add for: in this step.

The skeleton of a rules file is groups: → - name: → rules: → - alert: and expr:. For the forecast, give predict_linear a range vector and a future offset in seconds. For p99, give histogram_quantile the rate aggregated by le. The check is promtool check rules /root/obs-alert-symptom/rules/causes.yml.

Build a tool that counts how many times an alert would fire

Create /root/obs-alert-symptom/count.py. When called as python3 count.py '<경보 식>' <for 분> (the alert expression and the for duration in minutes), it scans the last 12 hours at 1-minute intervals and prints one line, episodes=<호출 수> minutes=<울린 분> (the number of pages and the minutes fired). When the minutes in which the condition is true continue for L minutes in a row, if L is for + 1 or more, count it as one page, and the time fired is L - for minutes. If for is not given, treat it as 0.

If you give start, end, and step to /api/v1/query_range, range data comes back. In minutes when the condition is false there is no sample at all, so build the set of timestamps in the response and count the continuous runs on a 1-minute grid. Use only the Python standard library (urllib.request, json, time).

How many times did the four cause alerts fire in 12 hours

Create /root/obs-alert-symptom/counts.tsv. It has four lines with no header, and each line has three tab-separated columns, <경보이름> <호출 수> <울린 분> (the alert name, the number of pages, and the minutes fired). Leave for at 0, and write the four alerts you wrote in the step 1 rules file by their names as they are. Get the values with the tool from step 2.

If you take the expressions from the rules file and pass them to the tool, you reduce the mistake of copying them by hand. How far the four numbers diverge is the heart of this step — one of them fires more than sixty times.

How much does one duration reduce the pages

Pick the alert that fired the most in step 3, and count again while changing for to 0, 5, and 15 minutes. In /root/obs-alert-symptom/for-effect.tsv, write three lines with no header, and each line is <for 분> <호출 수> <울린 분> (for minutes, number of pages, minutes fired).

Leave the condition expression as it is and change only the second argument. Also see what you lose in exchange for the pages decreasing — fewer minutes fired means detection is delayed by that much.

Write one symptom-based alert

In /root/obs-alert-symptom/rules/symptom.yml, write a symptom group and one alert, SymptomErrorRatio. The condition is that the 5xx response ratio over the last 5 minutes exceeds 1%, and attach for: 5m and a runbook_url annotation. After it passes promtool check rules, count the same expression with for at 0 and write one line, episodes=<수> minutes=<분> (number and minutes), to /root/obs-alert-symptom/symptom.txt.

For the ratio, divide the sum of the rate of 5xx by the sum of the total rate. The point is that this alert fires far less than the cause alerts. Put runbook_url under annotations.

How many minutes did it fire when users experienced nothing

Create /root/obs-alert-symptom/overlap.py. When called as python3 overlap.py '<원인 식>' '<증상 식>' (the cause expression and the symptom expression), it scans the last 12 hours at 1-minute intervals and prints one line, cause_minutes=<원인이 참인 분> outside_minutes=<그중 증상이 참이 아닌 분> (the minutes the cause is true, and of those the minutes the symptom is not true). Measure each of the four cause alerts with that tool and write four tab-separated three-column lines, <경보이름> <원인 분> <증상 밖 분> (alert name, cause minutes, minutes outside the symptom), to /root/obs-alert-symptom/falsepages.tsv.

Build the 'set of true minutes' of each of the two expressions and count the set difference. If you split half of the step 2 tool out into a function, you can use it as is. The larger the minutes outside the symptom, the longer the time that 'users were fine but only people were woken up'.

Split page, ticket, and dashboard by the numbers

Write five lines in /root/obs-alert-symptom/triage.tsv. Each line has three tab-separated columns, <경보이름> <page|ticket|dashboard> <근거> (alert name, the classification, and the grounds), and you write all four cause alerts and SymptomErrorRatio. There are two rules — the symptom alert must be page, and a cause alert whose minutes outside the symptom exceeded 60 in step 6 cannot be set as page. The grounds must include at least one number obtained in the previous steps and be at least 20 characters.

A cause alert whose minutes outside the symptom are close to 0 fires only at almost the same times as the symptom, so you can keep it as a page and save investigation time. Conversely, an alert that is true nearly half the time, if left as a page, makes people create filters.

Keep only the rules file for pages

In /root/obs-alert-symptom/rules/page.yml, include only the alerts you classified as page in step 7. The group name is page, and each alert must have for: of 5 minutes or more, severity: page in labels, and runbook_url in annotations. After it passes promtool check rules, count those alerts with each one's own for value and write the sum of the page counts as one line, pages_after=<수> (the number), to /root/obs-alert-symptom/after.txt.

If you read the alert name, expr, and for from the rules file and pass them to the step 2 tool, there is nothing to add by hand. The for value is a string like 5m, so you have to convert it to a number of minutes. Compare it with the sum of the page counts from step 3.