The Timeline Is Half the Report
One-line summary
A timeline is not a list of events but a tool that reveals what each interval means, and those intervals are exactly the time you should reduce next time.
Why this is needed
What the customer reads first in an outage report is not the root cause analysis but the timeline. There is a reason. A timeline shows at the same time "how long we did not know" and "how quickly we moved once we knew."
And these two numbers decide the improvement tasks for the next quarter.
How it works
A useful timeline has five times.
Change time — when something changed. A deployment, a configuration change, a traffic surge. Impact start time — when users actually began to experience failures. Awareness time — when we found out. When the alarm sounded or the customer reported it. Action time — when the mitigation went in. Recovery confirmation time — when we measured again with the same yardstick and confirmed normal.
And the four intervals between these five each have a name.
Change → impact: the incubation interval. If short, it surfaced immediately, and if long, it is the type that blows up only after certain conditions accumulate. The latter is far scarier.
Impact → awareness: detection delay. If this value is large, the problem is not the system but observation. If the situation where the customer tells you first repeats, trust is not restored no matter how well you fix the cause.
Awareness → action: response delay. It is short if there is a runbook and long if there is not.
Action → recovery confirmation: the verification interval. A report without this interval has said only up to "we think it is fixed."
What you see in the field
Let me point out one judgment that comes up often here. It is from where to where you count the impact duration.
If you count from the change time, it is longer than reality, and if you count only up to the action time, it is shorter than reality. The honest calculation from the user's point of view is from the impact start to the recovery confirmation. Users do not know when the deployment was, and they get out of the impact not at the moment the rollback command went in but at the moment things actually became normal.
And there is one more rule to keep when writing a timeline. Do not use people's names. Not "Mr. Kim deployed at 03:19" but "03:19 payment 2.7.0 deployed."
This is not morality but calculation. If a specific person is singled out in the customer report, at the next outage that person hides information. Then the next detection delay grows. A blameless report is an investment that buys speed for the next diagnosis.
Aligning the times is half of it
When you line up logs from several sources on a single line, the first things that trip you up are time zones and precision.
| Source | Common format | Trap |
|---|---|---|
| Application log | 2026-09-06 12:04:31 |
No time zone — you cannot tell whether it is server local or UTC |
| Kubernetes events | 2026-09-06T12:04:31Z |
Fixed to UTC. Nine hours apart from the app log |
| Cloud console | Browser local time | Looks different to each viewer |
| Alarm | Not the time it occurred but the evaluation time | 1–5 minutes later than reality |
Convert everything to UTC and state at the very top of the document that you did so. And mark that an alarm is the "detection time," not the "occurrence time." If you do not make this distinction, you mix up "the alert was late" and "the outage started late."
Match the precision too. When you line up a log in seconds and a log in milliseconds on the same line, round down to seconds. If you write precision that does not exist as if it does, causality looks reversed.
What to regard as causation
If you line things up in time order, you see the before-and-after relationship but not causation. You must confirm three things before you can call it causation.
- Time consistency — is the cause before the effect? Does the order hold even allowing for propagation delay?
- Scope consistency — is the cause's scope of impact the same as the symptom's scope? If only one Pod changed but everything died, there is another cause.
- Revert verification — did the symptom disappear when the cause was reverted? This is the strongest evidence.
If even one of the three does not match, write it as "seems related" and do not confirm it. Marking estimates and confirmations separately in the timeline document builds trust.
| 시각(UTC) | 사건 | 출처 | 확인 |
|---|---|---|---|
| 03:19:02 | payment 2.7.0 배포 시작 | Argo CD | 확인 |
| 03:19:40 | payment 파드 5xx 시작 | 앱 로그 | 확인 |
| 03:23:11 | 결제 실패율 경보 | Alertmanager | 확인(평가 시각) |
| 03:24 경 | 고객 문의 유입 | 지원 티켓 | 추정(티켓 시각은 접수 시각) |
| 03:31:05 | 2.6.4 로 롤백 | Argo CD | 확인 |
| 03:31:52 | 5xx 소멸 | 앱 로그 | 확인 — 되돌림으로 인과 성립 |
What you will do in the next lab
You pull out the five times each from the deployment history, the application log, and the alarm history, build a timeline document lined up in time order, and calculate the impact duration from the user's point of view.