TT Lab
Get started
Learn Learning paths Courses

In Front of an Unfamiliar System

Good Debugging Leaves No Trace

Continue in TT Lab

In one line

The commits of someone who narrowed it down precisely and fixed it in 30 minutes and of someone who wandered for three hours and fixed it by accident look exactly the same, so if you leave no record, neither skill nor organizational assets accumulate.

Why this was needed

There are three reasons debugging skill does not grow in proportion to years of experience.

First, almost nowhere is it taught explicitly. People learn languages and frameworks, but they learn the narrowing procedure by watching over someone's shoulder.

Second, the feedback is wired wrong. When the symptom disappears, the reward arrives. A state where you do not know why it disappeared is rewarded just the same. So the experience of fixing it by luck is mistaken for skill.

Third, good debugging leaves no trace. It is hard for an organization to recognize this ability, and what is not recognized is not cultivated either.

There is only one response to this. Leave a trace.

The narrowing procedure has names

To leave a record, you first need a procedure. If you give names to what experienced people do unconsciously, it becomes something that can be learned.

Binary search. Cut the layers a request passes through in half. If it is client → gateway → service → DB, first call the service directly from the gateway. If it works, the front part is the culprit; if not, the back part is. In one step the candidates are halved. With eight layers, you narrow to one in three steps.

Narrowing the difference (delta debugging). Use it when both a working case and a failing case exist. Erase the differences between the working request and the failing request one at a time, and find which difference, when it disappears, makes the symptom disappear too. One header, one field, or one time is what remains.

Aligning the timeline. Put the times of the deployment, the configuration change, and the traffic increase, together with the time the symptom started, on one line. Most outages start right after something changed. Knowing "since when" exactly is half of "because of what".

Flipping an assumption. If you have not narrowed it down after more than 30 minutes, it is very likely that one thing you believed was certain is wrong. "DNS obviously works", "the configuration was applied", "that version is the right one" — check each of these for real, once. The answer usually turns up here.

The format of the record — one page is enough

If you write at length, nobody writes. The format that actually gets maintained in practice is about this.

## 증상
결제 완료 화면에서 간헐적 500 (약 20%)

## 재현
for i in $(seq 20); do curl -s -o /dev/null -w '%{http_code}
'   -X POST https://stg/api/pay -d @fixtures/pay.json; done
→ 20회 중 4~6회 500

## 배제
- 인증: 401/403 이 로그에 0건
- 네트워크: 같은 요청을 게이트웨이 안에서 직접 → 같은 비율로 실패
- 데이터: 실패한 요청의 body 가 성공한 것과 바이트 단위로 동일

## 원인
결제 서비스 파드 3개 중 1개만 옛 설정(타임아웃 1초)으로 떠 있었다.
ConfigMap 을 바꾼 뒤 rollout restart 를 하지 않아 그 파드만 옛 값을 안고 있었다.

## 수리 전후
전: 20회 중 5회 실패 / 후: 40회 중 0회 실패 (같은 명령)

## 다음에 이걸 막는 것
- [ ] ConfigMap 해시를 파드 애노테이션에 넣어 변경 시 자동 재시작 (담당: 배포팀, 3/25)

"What prevents this next time" is the value of this document. A record that states only the cause is merely searchable when you meet the same outage for the second time, but a record written up to this point eliminates the second time.

Set a time box

The biggest cost of debugging is time passing without narrowing down. Decide in advance.

Elapsed What to do
15 minutes Write down in words what you have ruled out so far. While writing, you will see the layers you missed
30 minutes Check for real one assumption you believed was certain
45 minutes Call someone. Half of the time you find it yourself in the middle of explaining
60 minutes See whether there is a workaround. Finding the cause and restoring the service are different jobs

The last line is especially important. In outage response, recovery comes first and the cause comes later. Spending an hour finding the cause when you could recover in 5 minutes with a rollback is the wrong priority.

How it works

There are four kinds of records an FDE must leave.

The reproduction command. The minimal single line of command that causes the symptom. With it you can judge whether it was repaired against the same yardstick, and without it you can only say "it seems to be fixed".

The ruled-out list. The layers you checked and erased, and the evidence for them. One line is enough. For example, auth — 401/403 이 로그에 한 건도 없음 (the Korean text means "not a single 401/403 in the logs").

The before-and-after measurements. Before the repair 5 failures out of 20, after the repair 0 out of 20. What matters is that the two numbers were measured the same way.

The documents that remain after you leave. The success condition of an FDE is unusual. It counts as success only if it keeps running without me. If the system stops on the day the on-site residency ends, that was not a deployment but a rental.

What it looks like in the field

What customers reopen most often in a handoff document is not the architecture diagram but the outage runbook. A document that states, for each frequently occurring symptom, what to do in the first 30 minutes. A system without a runbook is a system that cannot be handed over.

And on the day you hand over the documents, handing over only the documents is not enough; completing the job means going as far as a rehearsal in which the customer's engineer handles one outage scenario directly by following the runbook. Having read and understood something is different from the hands moving.

More important than the format of the record is the timing. A handoff document must be built up from the middle of the project, not in the last week. A document written in a rush in the last week is, without exception, missing "what doesn't work". By then you are already used to those exceptions and they no longer feel special.

What you will do in the next lab

You actually leave the four kinds. And at the end, the grader plays the next person — it reruns the reproduction command you wrote in a clean environment that is not your shell. Only if the same number comes out there is the record really a record.