TT Lab
Get started
Learn Learning paths Courses

Observability

The first thing an investigation does is pin down the time

Continue in TT Lab

In one line

The first thing to do in an investigation is not to find the cause but to pin down the time. Only when the start and end are fixed can you count the scope of impact, and only then can you write, for each hypothesis, 'what must show up in the data'.

Why this matters

Once, four people gathered after receiving a payment error report. One mentioned "this morning's deployment", one "a traffic surge from a particular customer", and one "database connection exhaustion". All three were plausible and all three were stories that would take 30 minutes each to investigate. Two hours later the cause was a fourth candidate, and until then nobody had confirmed when the incident started from the data.

If the start time had been pinned down first, two of them would have been struck off on the spot. The deployment went out 40 minutes after the incident began, and the customer's traffic in question was the same as usual throughout the incident. Both facts would have come out of a single lookup. Just by swapping the order, two people's 30 minutes are saved.

What matters here is not memory but records. Human memory says "it seems to have started around 10", but data says exactly "the first moment the error ratio exceeded 0.05". The first line of an investigation must always be that condition expression and that time.

How it works

The order has six steps.

One, pin down the time. First fix the condition expression for what you count as the 'incident'. For example, "the 5xx ratio exceeds 5% in a 5-minute window". Then use a range query to find the time when that condition first became true and the time it was last true. At this point, do not write absolute times — time zones and daylight saving time ruin retrospectives. If you write it as "how many minutes ago from now", you can redraw the same interval later on anyone's screen.

Two, count the scope. How many failed, and which handler accounted for how much? A common trap here is cutting the window short. If you cover the incident period inadequately, the failure count drops out entirely. If you select only the minutes in which the condition was true and add up the increase in those minutes, you are not swayed by the window length.

Three, write the hypothesis and the prediction first. If you write only the hypothesis, you end up fitting the interpretation after looking at the data. If you first write "if this hypothesis is true, which metric should look how", you can discard the hypothesis when the data is not that shape.

Four, eliminate. Refutation is cheaper than confirmation. If it were a defect in one handler, only that handler's error rate should spike, but if all four handlers spike at the same ratio, that hypothesis is finished. If it were a traffic surge, the request rate should rise, but if the request rate in the incident period differs by 6% from the hour just before, that hypothesis is finished too.

Five, confirm what remains with another signal. Confirm the one left after elimination with a different metric from the initial condition expression. Looking at the same metric again is not confirmation but a repetition of the same statement. If you defined the incident by error rate, confirmation must be found in a signal that moves independently of the error rate, such as queue depth or saturation, and you must overlay, minute by minute, whether that signal moved in the same shape over the same interval.

Six, leave a machine-readable timeline. What the postmortem culture chapter of the Google SRE Book emphasizes is also 'what to leave'. For each event, one line with the relative time and the evidence metric. Sentences that people read cannot later be turned into statistics, but a table can.

What it looks like in the field

This order is what the effective troubleshooting chapter says over and over. However, the field most often missing from postmortems is not the cause but the detection delay. The difference between the time the impact started and the time the alert fired is almost never recorded. Yet this number decides what to fix next quarter. If you set a 10-minute persistence condition on a 5-minute window, detection takes 11 minutes, and if the incident was 20 minutes long, a person woke up only after half of it had passed.

Another is putting unrelated events on the same timeline. There is a separate period within the same 12 hours where only the tail latency spikes, and even though it does not overlap in time with the error incident, it often gets written into the retrospective document together. When putting events on a timeline, you must write the evidence metric on each line so that it remains later that the two are different events.

The third is the meeting flowing toward increasing the number of hypotheses. When there are six candidates, people scatter in six directions, each opening a different metric and looking at a different time axis. What an investigation needs is not the ability to add candidates but an order that eliminates cheaply. If you check the hypothesis with the most specific prediction first, one falls away per lookup, and once the remaining candidates are two or fewer, from then on several people can join without overlapping. Pasting in the queries used for elimination and their results as they are is worth more to the person who comes later than the conclusion.

The last is not leaving the investigation process itself. If you do not write down the hypotheses you eliminated and the evidence you eliminated them with, in the next incident you spend another 30 minutes on the same hypothesis. A record of elimination is worth as much as a record of the cause.

What you will do in the next lab

You investigate, in order, one incident contained in the 12 hours of the Pod's Prometheus. You fix the condition expression and pin down the start and end as relative times, select only the minutes in which the condition was true and count the failures and the per-handler distribution. After first writing three hypotheses and each one's prediction, you eliminate two with lookups and confirm the remaining one with another metric. Then you create a machine-readable timeline and postmortem data, and finally write a detection rule that catches it faster than now and verify, with the same 12 hours of data, how much faster it gets and whether false pages arise in exchange.