FDE Capstone: The Warehouse Got the Same Order Three Times
3 a.m., and the commands in the handoff doc were wrong
Goal
For the operations team that will take over, you build a chain of metric → alert condition → runbook entry → executable verification command, and the grader starts a fake service in various failure modes and checks through a drill whether that chain points to the exact entry.
Why it matters
A handoff document is correct only on the day it is written. If which conditions fire an alert and what to use to verify it when it fires are written down separately from the code, the document goes stale the moment the service changes, and that fact comes out when the on-call engineer types the command at dawn. So you keep the alert rules as data that a script reads, write the check lines of the runbook as commands that answer with an exit code, and repeatedly confirm with a checker and a drill that those commands actually run. Not reading a quiet dashboard (metrics that could not be scraped) as normal is part of the same chain.
The expected time is 75 minutes. Extend with +time before the default 60-minute session ends (up to 180 minutes). When the session ends, the files in /root/drill disappear, so keep any code you want to keep separately before it ends.
Materials
- Handoff memo (alert table and runbook rules):
/opt/lab/drill/brief.md - Fake service:
python3 /opt/lab/drill/fakesvc.py --port 9311 --mode normal(check the failure modes with--help) - The vendor's old runbook:
/opt/lab/drill/vendor-runbook.md - The grader does not use the server you started. It starts the fake service itself on an empty port in various modes and with various seeds and runs your scripts. If you memorize and hard-code the numbers provided, you will not pass.
Steps
- Start the fake service in normal mode and save the
/metricsresponse, unprocessed, to/root/drill/normal.prom. From the# TYPElines of the original, write이름 형식(the placeholders are the name and the type) one per line into/root/drill/families.txt. - Write
/root/drill/promtext.py.python3 promtext.py FILEprints a JSON array of objects{"name","labels","value","timestamp"}, one per sample. The value is a number, infinity and NaN are the strings"+Inf","-Inf", and"NaN", and if there is no timestamp, it is null. Skip comments and empty lines, and correctly read the commas and braces inside a label value and the\\,\", and\nescapes. If even one sample line is invalid, the exit code is 2. - Move the six alerts of the handoff memo into
/root/drill/alerts.json. The format is{"alerts": [{"alert","severity","runbook","rule"}]}, and rule.kind is one ofthreshold(metric, op, value),ratio(metric, denominator, op, value),expires_within(metric, seconds),scrape, andabsent(metrics). The grader judges your rules in various situations including boundary values. - Write
/root/drill/diagnose.py.python3 diagnose.py --url URL [--rules 경로](the placeholder is the path; the default is/root/drill/alerts.json) scrapesURL/metricsonce, judges with the rules, and prints{"target": URL, "firing": [{"alert","severity","runbook","labels"}]}. labels are the labels of that sample. The exit code is 0 if there are no alerts and 1 if there are. The thresholds are read from the rules file. - Fix
/root/drill/diagnose.pyso that it raises a failed scrape as an alert. On a connection failure, a non-200 HTTP response, a format error, or a missingorders_up, emit only oneOrdersTargetDown(labels{}) and exit with code 2. If a required metric is missing, emitOrdersMetricAbsent(labels{"metric": 이름}, where the placeholder is the metric name) for each missing metric. - Write the runbook entries for the six alerts in
/root/drill/runbook.md. Each entry is a## 런북이름heading (the placeholder is the runbook name) and five lines,- 경보:,- 확인:,- 판단:,- 조치:, and- 에스컬레이션:(the Korean keywords mean alert, check, judgment, action, and escalation). The check line is a one-line command wrapped in backticks that points to the service with$ORDERS_URL, and when run withbash -o pipefail -c, it must end right away with 0 if normal and a nonzero value if it is that failure. - Write
/root/drill/check_runbook.py.python3 check_runbook.py 런북.md --url URL(the placeholder is the runbook file) runs each entry's check command withbash -o pipefail -ctogether with theORDERS_URLenvironment variable within a 5-second limit, and prints[{"id","check","exit","status","reason"}]in runbook order. status isokif the exit code is 0 andbrokenotherwise. reason isok,timeout,not-found(127),exit N, orno-check(no check line). If even one is broken, the exit code is 1. Run it on the vendor runbook to find the broken lines. - Write the on-call drill script
/root/drill/oncall.sh.bash oncall.sh URLgets the alerts with diagnose.py, finds the runbook.md entry for each alert, runs the check command, and prints{"target": URL, "incidents": [{"alert","severity","runbook","labels","check_exit","confirmed","escalation"}]}. confirmed is true when the check command ended with a nonzero value (the failure is confirmed), and escalation is the runbook's escalation line as it is. If there are incidents, the exit code is 1, and otherwise 0.
Notes
- Saving the original:
curl -fsS http://127.0.0.1:9311/metrics -o 파일(the placeholder is the file) — if you process it through a pipe, the final newline or the comments can disappear. - When you start several services to compare, change the port (9312, 9313…). Take down a finished service by specifying the port too, like
pkill -f 'fakesvc.py --port 9312'. - Common mistake 1: splitting labels on commas. It breaks on
note="handoff \"v2\", see RB". - Common mistake 2: swallowing a scrape failure into an empty list. Then on a night the service is dead, the alerts come out as zero.
- Common mistake 3:
curl ... | grep -qin a check command. grep ends at the first match so curl receives SIGPIPE (exit code 23), and under pipefail it can look like a failure even when it is normal. Use a form that reads the input to the end, such asgrep -c ... >/dev/nullorjq -e. - Common mistake 4: applying the time limit only to bash. The curl after the pipe survives and the checker hangs with it. Start it in a new session and cut off the whole process group.
Scrape the metric text as it is
Save the /metrics response of the normal-mode fake service to /root/drill/normal.prom, and the list of 이름 형식 (the placeholders are the name and the type) from the TYPE lines to /root/drill/families.txt.
A # TYPE <이름> <형식> line (the placeholders are the name and the type) exists only once per metric family. For a histogram family, the name on the TYPE line is the reference, and the samples come out with _bucket, _sum, and _count attached. You can just pick, with awk, the lines whose first two fields are # and TYPE.
A parser that is not fooled by commas and quotes
Write /root/drill/promtext.py, which turns the Prometheus text exposition format into a JSON array of sample objects. If there is an invalid sample line, the exit code is 2.
You cannot cut the label part with split(','). Inside quotes, a comma and } are part of the value, and the character after a backslash is an escape. Build a small state machine (inside quotes? right after a backslash?) that reads one character at a time. Read the value with float(), but turn infinity and NaN, which JSON does not have, into strings.
The six alerts of the handoff memo as data
Move the alert table of /opt/lab/drill/brief.md into the /root/drill/alerts.json rules.
"Below" and "at or below", and "above" and "at or above", are different ops. The certificate metric is an expiry time, so use expires_within, and for the queue, use not the depth but the age of the oldest message. Write a failed scrape as scrape and the list of required metrics in absent's metrics.
Scrape once and judge with the rules
Write /root/drill/diagnose.py, which scrapes the metrics, judges with alerts.json, and outputs the fired alerts as JSON (no alerts 0 · alerts 1).
If you group the samples by name, you can pull out just what each kind of rule needs. For a ratio, you must pair the numerator and the denominator by the same labels. expires_within compares 'metric value − current time'. If you import promtext.py from the earlier step, you do not need to write the parser again.
Do not read a quiet dashboard as normal
Fix /root/drill/diagnose.py so that a failed scrape is raised as OrdersTargetDown with exit code 2 and a missing required metric as OrdersMetricAbsent.
urlopen throws URLError on a connection failure and HTTPError on a non-200 response. The parser's ParseError is a failed scrape too. If you turn these three into an empty sample list, every comparison becomes false and the alerts come out as zero — that itself must be an alert. absent is judged by 'is there not a single sample with this name'.
A runbook that answers with an exit code
Write the entries for the six alerts in /root/drill/runbook.md. The check commands must actually work as 0 when normal and nonzero for that failure.
The diagnostic endpoints (/debug/deps, /debug/disk, /debug/tls, /debug/queue) are JSON, so jq -e fits well. -e gives exit code 1 when the result is false or null. Add curl's -f (treat 4xx and 5xx as failures) and -m (time limit). target-down has to catch even a maintenance page that is 200, so you also check whether there is an orders_up sample in the response.
Run the vendor runbook and find the broken lines
Write /root/drill/check_runbook.py, which actually runs the runbook's check commands and judges them, and run it on /opt/lab/drill/vendor-runbook.md.
subprocess.run(timeout=…) kills only bash when the time passes, and the curl left after the pipe can hold the output pipe so that the wait never ends. Start it with Popen(start_new_session=True) and, when the time passes, cut off the whole group with os.killpg. Exit code 127 means 'command not found'.
An on-call drill from the alert to the escalation
Write /root/drill/oncall.sh, which runs diagnose.py → runbook.md → the check command → the escalation in one pass.
Exit codes 1 and 2 of diagnose.py are not errors but verdict results, so do not cut the script off with set -e. For the runbook parsing and the time-limited execution, you can import and reuse the functions of check_runbook.py from step 7. confirmed means 'did the check command confirm the failure', so it is true when the exit code is not 0.