TT Lab
Get started
Learn Learning paths Courses

FDE Capstone: The Warehouse Got the Same Order Three Times

Why handoff docs are wrong at 3 a.m.

Continue in TT Lab

In one line

The unit of an operations handoff is not a document but a chain. The alert conditions come from the metrics the service emits, each alert gets a runbook entry, and each entry needs a verification command that you can copy and run as it is. Which link of that chain is broken is revealed only by a drill.

Why this was needed

As a delivery nears its end, the FDE hands the service over to the customer's operations team. The usual handoff item is a single wiki page. The dashboard address, the alert list, a few lines saying "If something goes wrong, type this command". This document is correct on the day it is written. Three months later, at three in the morning, the on-call engineer gets a page and types that command, and the port has changed, the diagnostic path name is different, and that tool is not on the on-call engineer's laptop. On top of that, one command hangs with no response and eats up 30 seconds. The fact that the document is wrong comes out at the most expensive moment.

The on-call chapter of the Google SRE book puts the weight of this situation in numbers. For a user-facing service, a page response target of 5 minutes is common, and for a less urgent system it is 30 minutes. A single incident eats an average of 6 hours through root-cause analysis and postmortem, so they set an upper limit of two incidents per 12-hour on-call shift. And it warns that under stress people act on intuition and habit instead of deliberation, and that such reactions are easily wrong. That is why procedures are needed. What the on-call engineer at dawn needs is not a hero who will make the judgments for them, but prepared means of verification that let them avoid being wrong even while thinking less.

How it works

We look at the chain one link at a time.

1. Metrics. The Prometheus text exposition format is a line-oriented format. Lines are separated by newlines, the last line must also end with a newline, and empty lines are ignored. The # HELP and # TYPE lines give the description and the type (counter, gauge, histogram, summary, untyped), and there is only one TYPE line per metric name, which must come before the first sample. A sample has the form 이름{레이블="값"} 값 [타임스탬프] (the placeholders are the name, the label, the value, and the timestamp). Inside a label value, only three things are escaped, backslash, double quote, and newline, as \\, \", and \n. So a comma or a brace can appear inside a value, and a parser that splits on commas is silently wrong. A histogram is expanded into _bucket{le="..."}, _sum, and _count, and the le="+Inf" bucket must be present. Over HTTP it goes out as text/plain; version=0.0.4.

# HELP orders_queue_oldest_age_seconds Age of the oldest waiting message in seconds.
# TYPE orders_queue_oldest_age_seconds gauge
orders_queue_oldest_age_seconds{queue="fulfillment"} 2.4
orders_build_info{version="2.4.1",note="handoff \"v2\", see RB"} 1

2. Alert conditions. The monitoring chapter of the same SRE book says to separate symptoms from causes. What the user experiences (orders have not shipped for 10 minutes) is the symptom, and high CPU is a candidate cause. A page must be urgent, actionable, and in need of human judgment — a page where you end up mechanically typing the same command is something to automate or remove. By this standard, "queue depth over 1000" is a bad condition. At a lunch rush, even if the depth reaches the thousands, it drains within a few seconds. "The age of the oldest waiting message" is closer to the symptom users feel. Certificates are the same. Often the metric is not the time remaining but the expiry time, so you have to subtract the current time.

In a Prometheus alerting rule, the state in which expr is true must continue for the for period to move from pending to firing, and you put a description and a runbook link in annotations. There is no Prometheus server in this lab, so we carry the same idea over into small JSON rules and a script.

3. Quiet is not normal. A comparison expression becomes true or false only when the metric exists. If the metric is absent entirely, the result is empty and nothing fires. Prometheus has two devices for this. Each time it scrapes, it creates an up time series per target with 1 for success and 0 for failure, and the absent() function returns one element with the value 1 when its input is empty. That is why the documentation describes the purpose of this function as "for alerting on the absence of time series". The most dangerous thing is a dashboard that looks calm on a night when the collector is dead.

4. Runbook entries and executable verification. Each entry has four lines: check, judge, act, and escalate. The key is the check line. Do not leave the on-call engineer to interpret the shape of the output; make it answer with an exit code. 0 if normal, and nonzero if it is that failure. That way, the same answer comes out whether a person reads it at dawn or a script runs it in the drill. Here the pipe is a trap. In bash, the exit code of a pipe is by default that of the last command, so df 없는경로 | awk ... becomes 0 even when df fails (the placeholder is a path that does not exist). You have to run it with set -o pipefail to see the failure of the earlier command. Conversely, grep -q ends right at the first match, so the curl before it can receive SIGPIPE, so under pipefail you choose a form that reads the input to the end.

What it looks like in the field

If you run a runbook left by a vendor against a healthy service one line at a time, the kinds of breakage gather into a few groups. Commands that look at the on-call engineer's machine instead of the service (df -h /var/lib/orders looks at the disk of the on-call laptop), endpoints whose names have changed (404), old ports, tools missing from the on-call environment (in a container without ss, exit code 127), and commands that hang because they have no time limit. The last one makes the checker hang too. If you kill only bash with the timeout of subprocess, the curl after the pipe stays alive holding the output pipe, so you have to start it in a new session and cut off the whole process group.

This check is not done once. Every time the service changes, the handoff document goes stale too. So you include a "runbook checker" in the handoff package, and in a regular drill you turn on failure modes one at a time and see whether alert → runbook → verification command → escalation runs through to the end. This is what differs from a single record page (a reproducible runbook) or a postmortem. Those two record what has already happened, while a drill lets you live through in advance a night that has not happened yet.

What really matters in practice

What you will do in the next lab

You build the chain yourself while starting a fake orders service in various failure modes. You scrape the raw metric text to check the format, write a parser that also reads the escapes, turn the six alerts of the handoff memo into JSON rules, and judge them with a diagnostic script. You raise a failed scrape as an alert, and the grader runs the runbook verification commands in several modes to see whether they really run as 0 when normal and nonzero when failing. You find the broken lines of the vendor runbook with the checker, and finally build an on-call drill script that runs from the alert through the escalation in one pass.

References: Prometheus text exposition format · Prometheus alerting rules · Prometheus functions (absent) · Jobs and instances (the up series) · Google SRE book: Being On-Call · Google SRE book: Monitoring Distributed Systems