TT Lab
Get started
Learn Learning paths Courses

Observability

Two of the three hypotheses could have been ruled out from data in 30 minutes

Continue in TT Lab

Goal

You investigate, by procedure, one incident contained in the 12 hours of the Pod's Prometheus. You pin down the time as relative times, count the impact, write three hypotheses and their predictions first and then eliminate two, confirm the remaining one with another metric, leave a timeline and a postmortem, and finally write a detector that catches it faster and verify it with the same data.

Why it matters

The reason outage investigations take long is usually not that there is no data but that there is no order. If you state the hypothesis first and look at the data afterward, the interpretation is pulled toward the hypothesis. Conversely, if you first write 'if this hypothesis is true, what must show up', you can discard the hypothesis when the data is not that shape, and refutation is far cheaper than confirmation. What must come before that is pinning down the time — without a start and an end, something like 'the request rate in the incident period' cannot be computed at all. And all times are written as relative times. Absolute times ride on time zones, daylight saving time, and the log server's clock and ruin the retrospective. What to leave at the end is not only the cause — the hypotheses you eliminated, the evidence you eliminated them with, and how many minutes late detection was decide what to fix next quarter.

Steps

  1. Write the detection condition to /root/obs-incident-triage/onset.promql — the ratio that 5xx takes in a 5-minute window (you may include or omit the threshold comparison). Then sweep the last 12 hours with a range query at 1-minute intervals, find the period in which this ratio exceeded 0.05, and write four lines in /root/obs-incident-triage/01-onset.txt — start_minutes_ago= (how many minutes ago from now the time it first exceeded was, as an integer), end_minutes_ago= (the time it last exceeded), duration_min= (the number of minutes exceeded), and peak_ratio= (the maximum ratio in that period, to four decimal places). Do not write absolute times.
  2. Select only the minutes in which the condition was true and add up the increase(...[1m]) value of those minutes. Write three lines in /root/obs-incident-triage/02-impact.txt: failed_requests= (the 5xx count, an integer), total_requests= (the total number of requests, an integer), and failed_share= (the ratio of the two, to four decimal places), and in /root/obs-incident-triage/02-handlers.tsv, write the four handlers one per line with three tab-separated columns, <handler><탭><5xx 건수><탭><전체 5xx 중 비중> (handler, tab, 5xx count, tab, share of all 5xx; write the share to four decimal places). If you cut the window by time and count, the count drops out entirely at the boundary — select minutes and add them up.
  3. Write three lines in /root/obs-incident-triage/03-hypotheses.tsv. Each line has four tab-separated columns, <id><탭><가설><탭><예측><탭><검증에 쓸 PromQL> (id, tab, hypothesis, tab, prediction, tab, the PromQL to use for verification), and the ids are h1, h2, and h3. The hypotheses are h1 = a code defect in a particular handler, h2 = overload from a traffic surge, and h3 = a stall in a dependency downstream of the order queue. In the prediction column, write 'if this hypothesis is true, what must show up in the data' in at least 25 characters, and in the last column write a query that can confirm it. The query for h1 must split by handler, h2 must look at the rate of http_requests_total, and h3 must look at queue_depth.
  4. Pick the two that can be eliminated by the data among the three hypotheses and write them in two lines in /root/obs-incident-triage/04-refuted.tsv. Each line has four tab-separated columns, <id><탭><측정값><탭><판정><탭><근거> (id, tab, measured value, tab, verdict, tab, grounds), and the verdict is refuted. The measured value is fixed per id — for the handler hypothesis, it is the maximum minus the minimum among the 12-hour maximums of the per-handler 5xx ratio (5-minute window) (to four decimal places), and for the traffic hypothesis, it is the average requests per second in the incident period ÷ the average of the hour just before the incident started (excluding the last 5 minutes) (to four decimal places). The grounds must be at least 25 characters and include numbers.
  5. Write five lines in /root/obs-incident-triage/05-confirm.txt — hypothesis= (the id of the remaining hypothesis), queue_peak= (the maximum of queue_depth over the 12 hours, an integer), queue_baseline= (the median over the same 12 hours, an integer), overlap_min= (the number of minutes, among the minutes the queue depth exceeded 100, in which the error condition was also true, an integer), and evidence= (one line of at least 40 characters that includes numbers). Confirmation must be done with a different metric from the initial condition expression.
  6. Write four lines in /root/obs-incident-triage/06-timeline.tsv. Each line has three tab-separated columns, <event><탭><minutes_ago><탭><근거> (event, tab, minutes_ago, tab, grounds), and the events are the four impact_start, detect, impact_end, and latency_tail_start. detect is the time the alert currently in use (the 5-minute-window 5xx ratio staying above 0.10 for 10 minutes) would have fired, and latency_tail_start is the time p99 latency started to exceed 1 second (this is a different event from the error incident). minutes_ago is an integer, and in the grounds column write the name of the metric or query that produced the value, in at least 10 characters.
  7. Write five lines in /root/obs-incident-triage/07-postmortem.txt — time_to_detect_min= (how many minutes from the start of the impact to the alert, an integer), time_to_recover_min= (how many minutes from the start of the impact to the end of the impact, an integer), impact= (at least 40 characters, including numbers), why_late= (why detection was late, at least 60 characters), and next_change= (what to change next, at least 60 characters). The two numbers come straight from the step 6 timeline.
  8. Write a faster group and a FastErrorRatio alert in /root/obs-incident-triage/rules/faster.yml — it must have expr, for, labels.severity, and annotations.runbook_url, and for is in the form <정수>m (an integer followed by m). This rule must fire at least 3 minutes earlier than the current alert (5-minute window, 0.10, 10-minute persistence), and within the same 12 hours the number of minutes it fires outside the incident period must not exceed 10. Then write three lines in /root/obs-incident-triage/08-gain.txt, baseline_detect_min_ago=, new_detect_min_ago=, and gain_min= (the difference of the two values), as integers.

Notes

Pin down the start time from the data

Write the detection condition to /root/obs-incident-triage/onset.promql — the ratio that 5xx takes in a 5-minute window (you may include or omit the threshold comparison). Then sweep the last 12 hours with a range query at 1-minute intervals, find the period in which this ratio exceeded 0.05, and write four lines in /root/obs-incident-triage/01-onset.txt — start_minutes_ago= (how many minutes ago from now the time it first exceeded was, as an integer), end_minutes_ago= (the time it last exceeded), duration_min= (the number of minutes exceeded), and peak_ratio= (the maximum ratio in that period, to four decimal places). Do not write absolute times.

You throw a range query by attaching start, end, and step to /api/v1/query_range. With step=60, one value comes out per minute. Get now in seconds with date +%s and subtract 43200 to get 12 hours ago. You can turn the result JSON into a table with jq -r '.data.result[0].values[] | "\(.[0])\t\(.[1])"'.

Count the scope of impact — how many failed and where they concentrated

Select only the minutes in which the condition was true and add up the increase(...[1m]) value of those minutes. Write three lines in /root/obs-incident-triage/02-impact.txt: failed_requests= (the 5xx count, an integer), total_requests= (the total number of requests, an integer), and failed_share= (the ratio of the two, to four decimal places), and in /root/obs-incident-triage/02-handlers.tsv, write the four handlers one per line with three tab-separated columns, <handler><탭><5xx 건수><탭><전체 5xx 중 비중> (handler, tab, 5xx count, tab, share of all 5xx; write the share to four decimal places). If you cut the window by time and count, the count drops out entirely at the boundary — select minutes and add them up.

If you pull out only the timestamps whose value exceeds 0.05 from the ratio table you made in step 1, you can match those timestamps against the other lookup results and add them up (awk 'NR==FNR{m[$1]=1;next} ($1 in m){s+=$2}'). The handler names are /api/orders, /api/users, /api/search, and /healthz.

Write the prediction before the hypothesis

Write three lines in /root/obs-incident-triage/03-hypotheses.tsv. Each line has four tab-separated columns, <id><탭><가설><탭><예측><탭><검증에 쓸 PromQL> (id, tab, hypothesis, tab, prediction, tab, the PromQL to use for verification), and the ids are h1, h2, and h3. The hypotheses are h1 = a code defect in a particular handler, h2 = overload from a traffic surge, and h3 = a stall in a dependency downstream of the order queue. In the prediction column, write 'if this hypothesis is true, what must show up in the data' in at least 25 characters, and in the last column write a query that can confirm it. The query for h1 must split by handler, h2 must look at the rate of http_requests_total, and h3 must look at queue_depth.

A prediction comes out well if you flip it around: 'what would I see that lets me say this hypothesis is wrong?' For example, if it is a defect in one handler, the other handlers should be normal, and if all of them spike together, that hypothesis is eliminated. The three queries must actually return results — the grader throws them directly.

Eliminate two hypotheses with the data

Pick the two that can be eliminated by the data among the three hypotheses and write them in two lines in /root/obs-incident-triage/04-refuted.tsv. Each line has four tab-separated columns, <id><탭><측정값><탭><판정><탭><근거> (id, tab, measured value, tab, verdict, tab, grounds), and the verdict is refuted. The measured value is fixed per id — for the handler hypothesis, it is the maximum minus the minimum among the 12-hour maximums of the per-handler 5xx ratio (5-minute window) (to four decimal places), and for the traffic hypothesis, it is the average requests per second in the incident period ÷ the average of the hour just before the incident started (excluding the last 5 minutes) (to four decimal places). The grounds must be at least 25 characters and include numbers.

With four handlers there are four maximums. If only one handler was broken, these four diverge widely, and if the cause is common, they are almost the same. If the traffic ratio is near 1, it means 'it is not a surge'. You can get both numbers by sweeping a 12-hour range query, and for the incident period use the minutes you selected in step 1 as they are.

Confirm the remaining hypothesis with another signal

Write five lines in /root/obs-incident-triage/05-confirm.txt — hypothesis= (the id of the remaining hypothesis), queue_peak= (the maximum of queue_depth over the 12 hours, an integer), queue_baseline= (the median over the same 12 hours, an integer), overlap_min= (the number of minutes, among the minutes the queue depth exceeded 100, in which the error condition was also true, an integer), and evidence= (one line of at least 40 characters that includes numbers). Confirmation must be done with a different metric from the initial condition expression.

For the median, just extract the values, sort them, and pick the middle. The number of overlapping minutes is the size of the intersection of the set of minutes you selected in step 1 and the set of minutes the queue exceeded 100. If the two periods overlap almost completely, it is likely the same event, and that becomes a second piece of evidence independent of the error metric.

A machine-readable incident timeline

Write four lines in /root/obs-incident-triage/06-timeline.tsv. Each line has three tab-separated columns, <event><탭><minutes_ago><탭><근거> (event, tab, minutes_ago, tab, grounds), and the events are the four impact_start, detect, impact_end, and latency_tail_start. detect is the time the alert currently in use (the 5-minute-window 5xx ratio staying above 0.10 for 10 minutes) would have fired, and latency_tail_start is the time p99 latency started to exceed 1 second (this is a different event from the error incident). minutes_ago is an integer, and in the grounds column write the name of the metric or query that produced the value, in at least 10 characters.

A persistence condition (for) means 'after that many minutes of the condition being continuously true have accumulated' — if you count consecutive trues in the 1-minute-interval table, the minute that becomes the 11th is when the alert fires. You get p99 with histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))).

Postmortem — why was detection late

Write five lines in /root/obs-incident-triage/07-postmortem.txt — time_to_detect_min= (how many minutes from the start of the impact to the alert, an integer), time_to_recover_min= (how many minutes from the start of the impact to the end of the impact, an integer), impact= (at least 40 characters, including numbers), why_late= (why detection was late, at least 60 characters), and next_change= (what to change next, at least 60 characters). The two numbers come straight from the step 6 timeline.

Detection delay is a value created by the alert's window length and persistence condition — a 5-minute window creates delay by itself, and for adds to it. If you write in next_change= the direction of the rule you will actually build in the next step, it connects. Recovery time is how long the incident lasted, not what people did.

Write a detector that catches it faster and verify it with the same data

Write a faster group and a FastErrorRatio alert in /root/obs-incident-triage/rules/faster.yml — it must have expr, for, labels.severity, and annotations.runbook_url, and for is in the form <정수>m (an integer followed by m). This rule must fire at least 3 minutes earlier than the current alert (5-minute window, 0.10, 10-minute persistence), and within the same 12 hours the number of minutes it fires outside the incident period must not exceed 10. Then write three lines in /root/obs-incident-triage/08-gain.txt, baseline_detect_min_ago=, new_detect_min_ago=, and gain_min= (the difference of the two values), as integers.

Shortening the window makes it faster but also makes it noisier — so you must also check 'are there no false pages' with the same data. Look at the normal-time error ratio in the step 1 table, and if you choose a threshold sufficiently higher than that, it stays quiet even with a short window. Check the format first with promtool check rules <파일> (the placeholder is the file).