In Front of an Unfamiliar System
If It Is Not Falsifiable It Is a Guess
In one line
Debugging is not producing one possible cause but erasing the possible causes until only one is left, and to erase them, what you would see if a hypothesis were wrong has to be decided beforehand.
Why this was needed
"It's slow sometimes" cannot be diagnosed. All that can be diagnosed is a measured value. So after you take in the customer's words, your first task is not to find the cause but to pin down the symptom with a single command.
# 20회 반복해 실패율을 잰다
for i in $(seq 1 20); do
curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8001/quote
done | sort | uniq -c
If 5 out of 20 fail, the outage turns from a story you heard into a number. And this command is reused in every later step as the yardstick that judges whether it has been repaired. To say you fixed it, you have to measure again with the same yardstick and get 0.
Not being able to reproduce it is also information. It means it happens only for a specific user, in a specific time window, or on a specific path, and the question narrows by that much.
How it works
Once you have a number in hand, the next thing is a hypothesis. But what most people here call a hypothesis is not one.
"It seems to be a network problem" is not a hypothesis. Because what you would see if it were wrong is not decided, this sentence survives whatever result comes out. A guess that survives does not reduce the candidates.
A usable hypothesis takes three lines.
가설: 요청이 게이트웨이와 애플리케이션 사이에서 끊긴다
맞다면: 게이트웨이 로그에 5초 타임아웃, 앱 로그에 해당 요청 아이디 없음
틀리면: 앱 로그에 요청이 들어와 있고 처리 중 실패한 흔적이 있음
The difference is in the last two lines. What to check is specified, and if it comes out different from what was expected, the hypothesis dies. Only a hypothesis that can die reduces the candidates.
Writing the three lines takes 1 minute. That 1 minute looks like a waste, but if you do not write it down, 30 minutes later you find yourself rechecking what you have already ruled out. In an outage where several people are involved, the effect is greater. If the list of ruled-out items is not shared, three people check the same thing three times.
What it looks like in the field
The most trustworthy procedure for narrowing a hypothesis is to cut the layers in half. Write the layers a request passes through in order — client, DNS, load balancer, gateway, application, cache, database, external integration — then pick exactly the middle one and ask, "is everything up to here normal?"
If you sweep eight layers in order, you have to look four times on average, and if you cut in half, you are done within three. If the layers increase to twenty, the gap becomes ten versus five. In a real outage, where one check takes several minutes, this difference is the recovery time.
The key here is to pick a point where you can judge whether it is normal. If you cut at a point where judging is impossible, you cannot erase the half. The purpose of splitting layers is not to hit the right answer but to establish the places you do not need to look.
There is one exception. While the service is down right now, mitigation comes before diagnosis. Stop the bleeding by rolling back or redirecting traffic, and narrow down afterwards. Writing the hypothesis neatly is something to do when users are not waiting.
The format for writing a hypothesis as a sentence
The difference between a guess in your head and a verifiable hypothesis is whether you can write it as a sentence.
❌ "DB 가 느린 것 같다"
→ 무엇을 확인하면 아닌 줄 알 수 있나? 정해져 있지 않다
✅ "결제 API 의 p95 지연 3초 중 2초 이상이 orders 테이블 조회에서 난다."
확인 방법: 그 구간에 타이머를 넣고 100건을 재본다
반증 조건: 조회가 200ms 이하면 이 가설은 틀렸다
Writing the falsification condition first is the key. Without it, whatever result comes out ends up in "it could be", and the investigation never ends.
Change only one thing at a time
If you change several things at once, you do not know which one worked. And usually one gets better and one gets worse, and they cancel each other out.
❌ 인덱스 추가 + 커넥션 풀 확대 + 캐시 도입 → 20% 좋아짐
무엇 때문인가? 셋 중 하나는 오히려 해로웠을 수도 있다
✅ 인덱스 추가만 → 측정 → 되돌림 → 커넥션 풀만 → 측정 …
If you have no time and must put in several together, at least so that they can be reverted, keep each in a separate commit. Later you can take them out one at a time and check.
Making the measurement trustworthy
Measure three times under the same conditions and look at the width of the wobble. An improvement smaller than that is not an improvement.
전: 2.9s, 3.1s, 3.0s → 폭 0.2s
후: 2.8s, 2.9s, 2.85s → 폭 0.1s, 평균 0.15s 개선
→ 흔들림과 개선폭이 비슷하다. 아직 결론을 낼 수 없다
And check that the measurement itself does not change the target. If you turn on logging to measure, you make the logging slow things down, and if you attach a profiler, the profiler's cost gets mixed in. If possible, measure under the same conditions as production, and if you cannot, record that difference.
Conditions for ending the investigation
Testing a hypothesis has to have an end. You end it on any one of three.
- You identified the cause and confirmed it by reverting — the best ending.
- The symptom disappeared and you know why it disappeared — when it is hard to confirm by reverting.
- You checked this far and the cost of digging further is greater than the benefit — this too is a legitimate ending. But you must write down how far you looked, so that the next person picks up from there.
If you do not accept number 3, investigations drag on for days. If there is a mitigation and recurrence is rare, stopping is often the right call.
What you will do in the next lab
Against a quote API that received a "it doesn't work sometimes" report, you pin down the failure rate as a number, write a three-line hypothesis as a document, record the layers you ruled out along with the evidence, and then measure again with the same yardstick to check whether the ratio is stable.