Two of the three hypotheses could have been ruled out from data in 30 minutes
Goal
You investigate, by procedure, one incident contained in the 12 hours of the Pod's Prometheus. You pin down the time as relative times, count the impact, write three hypotheses and their predictions first and then eliminate two, confirm the remaining one with another metric, leave a timeline and a postmortem, and finally write a detector that catches it faster and verify it with the same data.
Why it matters
The reason outage investigations take long is usually not that there is no data but that there is no order. If you state the hypothesis first and look at the data afterward, the interpretation is pulled toward the hypothesis. Conversely, if you first write 'if this hypothesis is true, what must show up', you can discard the hypothesis when the data is not that shape, and refutation is far cheaper than confirmation. What must come before that is pinning down the time — without a start and an end, something like 'the request rate in the incident period' cannot be computed at all. And all times are written as relative times. Absolute times ride on time zones, daylight saving time, and the log server's clock and ruin the retrospective. What to leave at the end is not only the cause — the hypotheses you eliminated, the evidence you eliminated them with, and how many minutes late detection was decide what to fix next quarter.
Steps
- Write the detection condition to
/root/obs-incident-triage/onset.promql— the ratio that 5xx takes in a 5-minute window (you may include or omit the threshold comparison). Then sweep the last 12 hours with a range query at 1-minute intervals, find the period in which this ratio exceeded 0.05, and write four lines in/root/obs-incident-triage/01-onset.txt—start_minutes_ago=(how many minutes ago from now the time it first exceeded was, as an integer),end_minutes_ago=(the time it last exceeded),duration_min=(the number of minutes exceeded), andpeak_ratio=(the maximum ratio in that period, to four decimal places). Do not write absolute times. - Select only the minutes in which the condition was true and add up the
increase(...[1m])value of those minutes. Write three lines in/root/obs-incident-triage/02-impact.txt:failed_requests=(the 5xx count, an integer),total_requests=(the total number of requests, an integer), andfailed_share=(the ratio of the two, to four decimal places), and in/root/obs-incident-triage/02-handlers.tsv, write the four handlers one per line with three tab-separated columns,<handler><탭><5xx 건수><탭><전체 5xx 중 비중>(handler, tab, 5xx count, tab, share of all 5xx; write the share to four decimal places). If you cut the window by time and count, the count drops out entirely at the boundary — select minutes and add them up. - Write three lines in
/root/obs-incident-triage/03-hypotheses.tsv. Each line has four tab-separated columns,<id><탭><가설><탭><예측><탭><검증에 쓸 PromQL>(id, tab, hypothesis, tab, prediction, tab, the PromQL to use for verification), and the ids areh1,h2, andh3. The hypotheses are h1 = a code defect in a particular handler, h2 = overload from a traffic surge, and h3 = a stall in a dependency downstream of the order queue. In the prediction column, write 'if this hypothesis is true, what must show up in the data' in at least 25 characters, and in the last column write a query that can confirm it. The query for h1 must split byhandler, h2 must look at therateofhttp_requests_total, and h3 must look atqueue_depth. - Pick the two that can be eliminated by the data among the three hypotheses and write them in two lines in
/root/obs-incident-triage/04-refuted.tsv. Each line has four tab-separated columns,<id><탭><측정값><탭><판정><탭><근거>(id, tab, measured value, tab, verdict, tab, grounds), and the verdict isrefuted. The measured value is fixed per id — for the handler hypothesis, it is the maximum minus the minimum among the 12-hour maximums of the per-handler 5xx ratio (5-minute window) (to four decimal places), and for the traffic hypothesis, it is the average requests per second in the incident period ÷ the average of the hour just before the incident started (excluding the last 5 minutes) (to four decimal places). The grounds must be at least 25 characters and include numbers. - Write five lines in
/root/obs-incident-triage/05-confirm.txt—hypothesis=(the id of the remaining hypothesis),queue_peak=(the maximum ofqueue_depthover the 12 hours, an integer),queue_baseline=(the median over the same 12 hours, an integer),overlap_min=(the number of minutes, among the minutes the queue depth exceeded 100, in which the error condition was also true, an integer), andevidence=(one line of at least 40 characters that includes numbers). Confirmation must be done with a different metric from the initial condition expression. - Write four lines in
/root/obs-incident-triage/06-timeline.tsv. Each line has three tab-separated columns,<event><탭><minutes_ago><탭><근거>(event, tab, minutes_ago, tab, grounds), and the events are the fourimpact_start,detect,impact_end, andlatency_tail_start.detectis the time the alert currently in use (the 5-minute-window 5xx ratio staying above 0.10 for 10 minutes) would have fired, andlatency_tail_startis the time p99 latency started to exceed 1 second (this is a different event from the error incident). minutes_ago is an integer, and in the grounds column write the name of the metric or query that produced the value, in at least 10 characters. - Write five lines in
/root/obs-incident-triage/07-postmortem.txt—time_to_detect_min=(how many minutes from the start of the impact to the alert, an integer),time_to_recover_min=(how many minutes from the start of the impact to the end of the impact, an integer),impact=(at least 40 characters, including numbers),why_late=(why detection was late, at least 60 characters), andnext_change=(what to change next, at least 60 characters). The two numbers come straight from the step 6 timeline. - Write a
fastergroup and aFastErrorRatioalert in/root/obs-incident-triage/rules/faster.yml— it must haveexpr,for,labels.severity, andannotations.runbook_url, andforis in the form<정수>m(an integer followed by m). This rule must fire at least 3 minutes earlier than the current alert (5-minute window, 0.10, 10-minute persistence), and within the same 12 hours the number of minutes it fires outside the incident period must not exceed 10. Then write three lines in/root/obs-incident-triage/08-gain.txt,baseline_detect_min_ago=,new_detect_min_ago=, andgain_min=(the difference of the two values), as integers.
Notes
- The working directory is
/root/obs-incident-triage. If it does not exist, create it first. - A single query is thrown with
promq "<PromQL>"(the placeholder is the query). This lab uses range queries a lot, so get comfortable with the formcurl -sG --data-urlencode "query=..." --data-urlencode "start=..." --data-urlencode "end=..." --data-urlencode "step=60" http://127.0.0.1:9090/api/v1/query_range. - The Pod's Prometheus has 12 hours preloaded and one incident is contained in it. There is one more separate period where only the tail latency spikes, and it is a different event from this incident.
- Do not use absolute times. Write every time as 'how many minutes ago from now'. The grader looks at it that way too.
- Common mistake: cutting the incident period by time and counting. If the boundary is off, the failure count drops out entirely — select the minutes in which the condition was true and add them up.
- Common mistake: re-checking the remaining hypothesis with the condition expression you first wrote. That is not confirmation but a repetition of the same statement.
- Postmortem Culture (SRE Book, Chapter 15) · Effective Troubleshooting (SRE Book, Chapter 12) · Managing Incidents (SRE Book, Chapter 14) · Alerting rules · Query API (range queries)
Pin down the start time from the data
Write the detection condition to /root/obs-incident-triage/onset.promql — the ratio that 5xx takes in a 5-minute window (you may include or omit the threshold comparison). Then sweep the last 12 hours with a range query at 1-minute intervals, find the period in which this ratio exceeded 0.05, and write four lines in /root/obs-incident-triage/01-onset.txt — start_minutes_ago= (how many minutes ago from now the time it first exceeded was, as an integer), end_minutes_ago= (the time it last exceeded), duration_min= (the number of minutes exceeded), and peak_ratio= (the maximum ratio in that period, to four decimal places). Do not write absolute times.
You throw a range query by attaching start, end, and step to /api/v1/query_range. With step=60, one value comes out per minute. Get now in seconds with date +%s and subtract 43200 to get 12 hours ago. You can turn the result JSON into a table with jq -r '.data.result[0].values[] | "\(.[0])\t\(.[1])"'.
Count the scope of impact — how many failed and where they concentrated
Select only the minutes in which the condition was true and add up the increase(...[1m]) value of those minutes. Write three lines in /root/obs-incident-triage/02-impact.txt: failed_requests= (the 5xx count, an integer), total_requests= (the total number of requests, an integer), and failed_share= (the ratio of the two, to four decimal places), and in /root/obs-incident-triage/02-handlers.tsv, write the four handlers one per line with three tab-separated columns, <handler><탭><5xx 건수><탭><전체 5xx 중 비중> (handler, tab, 5xx count, tab, share of all 5xx; write the share to four decimal places). If you cut the window by time and count, the count drops out entirely at the boundary — select minutes and add them up.
If you pull out only the timestamps whose value exceeds 0.05 from the ratio table you made in step 1, you can match those timestamps against the other lookup results and add them up (awk 'NR==FNR{m[$1]=1;next} ($1 in m){s+=$2}'). The handler names are /api/orders, /api/users, /api/search, and /healthz.
Write the prediction before the hypothesis
Write three lines in /root/obs-incident-triage/03-hypotheses.tsv. Each line has four tab-separated columns, <id><탭><가설><탭><예측><탭><검증에 쓸 PromQL> (id, tab, hypothesis, tab, prediction, tab, the PromQL to use for verification), and the ids are h1, h2, and h3. The hypotheses are h1 = a code defect in a particular handler, h2 = overload from a traffic surge, and h3 = a stall in a dependency downstream of the order queue. In the prediction column, write 'if this hypothesis is true, what must show up in the data' in at least 25 characters, and in the last column write a query that can confirm it. The query for h1 must split by handler, h2 must look at the rate of http_requests_total, and h3 must look at queue_depth.
A prediction comes out well if you flip it around: 'what would I see that lets me say this hypothesis is wrong?' For example, if it is a defect in one handler, the other handlers should be normal, and if all of them spike together, that hypothesis is eliminated. The three queries must actually return results — the grader throws them directly.
Eliminate two hypotheses with the data
Pick the two that can be eliminated by the data among the three hypotheses and write them in two lines in /root/obs-incident-triage/04-refuted.tsv. Each line has four tab-separated columns, <id><탭><측정값><탭><판정><탭><근거> (id, tab, measured value, tab, verdict, tab, grounds), and the verdict is refuted. The measured value is fixed per id — for the handler hypothesis, it is the maximum minus the minimum among the 12-hour maximums of the per-handler 5xx ratio (5-minute window) (to four decimal places), and for the traffic hypothesis, it is the average requests per second in the incident period ÷ the average of the hour just before the incident started (excluding the last 5 minutes) (to four decimal places). The grounds must be at least 25 characters and include numbers.
With four handlers there are four maximums. If only one handler was broken, these four diverge widely, and if the cause is common, they are almost the same. If the traffic ratio is near 1, it means 'it is not a surge'. You can get both numbers by sweeping a 12-hour range query, and for the incident period use the minutes you selected in step 1 as they are.
Confirm the remaining hypothesis with another signal
Write five lines in /root/obs-incident-triage/05-confirm.txt — hypothesis= (the id of the remaining hypothesis), queue_peak= (the maximum of queue_depth over the 12 hours, an integer), queue_baseline= (the median over the same 12 hours, an integer), overlap_min= (the number of minutes, among the minutes the queue depth exceeded 100, in which the error condition was also true, an integer), and evidence= (one line of at least 40 characters that includes numbers). Confirmation must be done with a different metric from the initial condition expression.
For the median, just extract the values, sort them, and pick the middle. The number of overlapping minutes is the size of the intersection of the set of minutes you selected in step 1 and the set of minutes the queue exceeded 100. If the two periods overlap almost completely, it is likely the same event, and that becomes a second piece of evidence independent of the error metric.
A machine-readable incident timeline
Write four lines in /root/obs-incident-triage/06-timeline.tsv. Each line has three tab-separated columns, <event><탭><minutes_ago><탭><근거> (event, tab, minutes_ago, tab, grounds), and the events are the four impact_start, detect, impact_end, and latency_tail_start. detect is the time the alert currently in use (the 5-minute-window 5xx ratio staying above 0.10 for 10 minutes) would have fired, and latency_tail_start is the time p99 latency started to exceed 1 second (this is a different event from the error incident). minutes_ago is an integer, and in the grounds column write the name of the metric or query that produced the value, in at least 10 characters.
A persistence condition (for) means 'after that many minutes of the condition being continuously true have accumulated' — if you count consecutive trues in the 1-minute-interval table, the minute that becomes the 11th is when the alert fires. You get p99 with histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))).
Postmortem — why was detection late
Write five lines in /root/obs-incident-triage/07-postmortem.txt — time_to_detect_min= (how many minutes from the start of the impact to the alert, an integer), time_to_recover_min= (how many minutes from the start of the impact to the end of the impact, an integer), impact= (at least 40 characters, including numbers), why_late= (why detection was late, at least 60 characters), and next_change= (what to change next, at least 60 characters). The two numbers come straight from the step 6 timeline.
Detection delay is a value created by the alert's window length and persistence condition — a 5-minute window creates delay by itself, and for adds to it. If you write in next_change= the direction of the rule you will actually build in the next step, it connects. Recovery time is how long the incident lasted, not what people did.
Write a detector that catches it faster and verify it with the same data
Write a faster group and a FastErrorRatio alert in /root/obs-incident-triage/rules/faster.yml — it must have expr, for, labels.severity, and annotations.runbook_url, and for is in the form <정수>m (an integer followed by m). This rule must fire at least 3 minutes earlier than the current alert (5-minute window, 0.10, 10-minute persistence), and within the same 12 hours the number of minutes it fires outside the incident period must not exceed 10. Then write three lines in /root/obs-incident-triage/08-gain.txt, baseline_detect_min_ago=, new_detect_min_ago=, and gain_min= (the difference of the two values), as integers.
Shortening the window makes it faster but also makes it noisier — so you must also check 'are there no false pages' with the same data. Look at the normal-time error ratio in the step 1 table, and if you choose a threshold sufficiently higher than that, it stays quiet even with a short window. Check the format first with promtool check rules <파일> (the placeholder is the file).