TT Lab
Get started
Learn Learning paths Courses

FDE Capstone: The Warehouse Got the Same Order Three Times

3 a.m., and the commands in the handoff doc were wrong

Continue in TT Lab

Goal

For the operations team that will take over, you build a chain of metric → alert condition → runbook entry → executable verification command, and the grader starts a fake service in various failure modes and checks through a drill whether that chain points to the exact entry.

Why it matters

A handoff document is correct only on the day it is written. If which conditions fire an alert and what to use to verify it when it fires are written down separately from the code, the document goes stale the moment the service changes, and that fact comes out when the on-call engineer types the command at dawn. So you keep the alert rules as data that a script reads, write the check lines of the runbook as commands that answer with an exit code, and repeatedly confirm with a checker and a drill that those commands actually run. Not reading a quiet dashboard (metrics that could not be scraped) as normal is part of the same chain.

The expected time is 75 minutes. Extend with +time before the default 60-minute session ends (up to 180 minutes). When the session ends, the files in /root/drill disappear, so keep any code you want to keep separately before it ends.

Materials

Steps

  1. Start the fake service in normal mode and save the /metrics response, unprocessed, to /root/drill/normal.prom. From the # TYPE lines of the original, write 이름 형식 (the placeholders are the name and the type) one per line into /root/drill/families.txt.
  2. Write /root/drill/promtext.py. python3 promtext.py FILE prints a JSON array of objects {"name","labels","value","timestamp"}, one per sample. The value is a number, infinity and NaN are the strings "+Inf", "-Inf", and "NaN", and if there is no timestamp, it is null. Skip comments and empty lines, and correctly read the commas and braces inside a label value and the \\, \", and \n escapes. If even one sample line is invalid, the exit code is 2.
  3. Move the six alerts of the handoff memo into /root/drill/alerts.json. The format is {"alerts": [{"alert","severity","runbook","rule"}]}, and rule.kind is one of threshold (metric, op, value), ratio (metric, denominator, op, value), expires_within (metric, seconds), scrape, and absent (metrics). The grader judges your rules in various situations including boundary values.
  4. Write /root/drill/diagnose.py. python3 diagnose.py --url URL [--rules 경로] (the placeholder is the path; the default is /root/drill/alerts.json) scrapes URL/metrics once, judges with the rules, and prints {"target": URL, "firing": [{"alert","severity","runbook","labels"}]}. labels are the labels of that sample. The exit code is 0 if there are no alerts and 1 if there are. The thresholds are read from the rules file.
  5. Fix /root/drill/diagnose.py so that it raises a failed scrape as an alert. On a connection failure, a non-200 HTTP response, a format error, or a missing orders_up, emit only one OrdersTargetDown (labels {}) and exit with code 2. If a required metric is missing, emit OrdersMetricAbsent (labels {"metric": 이름}, where the placeholder is the metric name) for each missing metric.
  6. Write the runbook entries for the six alerts in /root/drill/runbook.md. Each entry is a ## 런북이름 heading (the placeholder is the runbook name) and five lines, - 경보:, - 확인:, - 판단:, - 조치:, and - 에스컬레이션: (the Korean keywords mean alert, check, judgment, action, and escalation). The check line is a one-line command wrapped in backticks that points to the service with $ORDERS_URL, and when run with bash -o pipefail -c, it must end right away with 0 if normal and a nonzero value if it is that failure.
  7. Write /root/drill/check_runbook.py. python3 check_runbook.py 런북.md --url URL (the placeholder is the runbook file) runs each entry's check command with bash -o pipefail -c together with the ORDERS_URL environment variable within a 5-second limit, and prints [{"id","check","exit","status","reason"}] in runbook order. status is ok if the exit code is 0 and broken otherwise. reason is ok, timeout, not-found (127), exit N, or no-check (no check line). If even one is broken, the exit code is 1. Run it on the vendor runbook to find the broken lines.
  8. Write the on-call drill script /root/drill/oncall.sh. bash oncall.sh URL gets the alerts with diagnose.py, finds the runbook.md entry for each alert, runs the check command, and prints {"target": URL, "incidents": [{"alert","severity","runbook","labels","check_exit","confirmed","escalation"}]}. confirmed is true when the check command ended with a nonzero value (the failure is confirmed), and escalation is the runbook's escalation line as it is. If there are incidents, the exit code is 1, and otherwise 0.

Notes

Scrape the metric text as it is

Save the /metrics response of the normal-mode fake service to /root/drill/normal.prom, and the list of 이름 형식 (the placeholders are the name and the type) from the TYPE lines to /root/drill/families.txt.

A # TYPE <이름> <형식> line (the placeholders are the name and the type) exists only once per metric family. For a histogram family, the name on the TYPE line is the reference, and the samples come out with _bucket, _sum, and _count attached. You can just pick, with awk, the lines whose first two fields are # and TYPE.

A parser that is not fooled by commas and quotes

Write /root/drill/promtext.py, which turns the Prometheus text exposition format into a JSON array of sample objects. If there is an invalid sample line, the exit code is 2.

You cannot cut the label part with split(','). Inside quotes, a comma and } are part of the value, and the character after a backslash is an escape. Build a small state machine (inside quotes? right after a backslash?) that reads one character at a time. Read the value with float(), but turn infinity and NaN, which JSON does not have, into strings.

The six alerts of the handoff memo as data

Move the alert table of /opt/lab/drill/brief.md into the /root/drill/alerts.json rules.

"Below" and "at or below", and "above" and "at or above", are different ops. The certificate metric is an expiry time, so use expires_within, and for the queue, use not the depth but the age of the oldest message. Write a failed scrape as scrape and the list of required metrics in absent's metrics.

Scrape once and judge with the rules

Write /root/drill/diagnose.py, which scrapes the metrics, judges with alerts.json, and outputs the fired alerts as JSON (no alerts 0 · alerts 1).

If you group the samples by name, you can pull out just what each kind of rule needs. For a ratio, you must pair the numerator and the denominator by the same labels. expires_within compares 'metric value − current time'. If you import promtext.py from the earlier step, you do not need to write the parser again.

Do not read a quiet dashboard as normal

Fix /root/drill/diagnose.py so that a failed scrape is raised as OrdersTargetDown with exit code 2 and a missing required metric as OrdersMetricAbsent.

urlopen throws URLError on a connection failure and HTTPError on a non-200 response. The parser's ParseError is a failed scrape too. If you turn these three into an empty sample list, every comparison becomes false and the alerts come out as zero — that itself must be an alert. absent is judged by 'is there not a single sample with this name'.

A runbook that answers with an exit code

Write the entries for the six alerts in /root/drill/runbook.md. The check commands must actually work as 0 when normal and nonzero for that failure.

The diagnostic endpoints (/debug/deps, /debug/disk, /debug/tls, /debug/queue) are JSON, so jq -e fits well. -e gives exit code 1 when the result is false or null. Add curl's -f (treat 4xx and 5xx as failures) and -m (time limit). target-down has to catch even a maintenance page that is 200, so you also check whether there is an orders_up sample in the response.

Run the vendor runbook and find the broken lines

Write /root/drill/check_runbook.py, which actually runs the runbook's check commands and judges them, and run it on /opt/lab/drill/vendor-runbook.md.

subprocess.run(timeout=…) kills only bash when the time passes, and the curl left after the pipe can hold the output pipe so that the wait never ends. Start it with Popen(start_new_session=True) and, when the time passes, cut off the whole group with os.killpg. Exit code 127 means 'command not found'.

An on-call drill from the alert to the escalation

Write /root/drill/oncall.sh, which runs diagnose.py → runbook.md → the check command → the escalation in one pass.

Exit codes 1 and 2 of diagnose.py are not errors but verdict results, so do not cut the script off with set -e. For the runbook parsing and the time-limited execution, you can import and reuse the functions of check_runbook.py from step 7. confirmed means 'did the check command confirm the failure', so it is true when the exit code is not 0.