TT Lab
Get started
Learn Learning paths Courses

Integration and Deployment

What You Can Promise Is the Next Update, Not the Fix Time

Continue in TT Lab

Summary

An FDE's incident response ends not in the system but in the report, and even if it is technically fixed perfectly, if the report is late or vague, only anxiety remains in the customer's memory.

Why this was needed

Of the eight technical domains, only customer communication is not technical. But it is the channel that delivers the other seven to the customer, so if it is blocked, the project fails even if the diagnosis was right.

The most common failure comes from expectation management. And that failure usually arises because the unit of the promise is wrong.

Saying "we will solve it by this afternoon" at a stage where you do not know the cause is not a promise but a gamble. What you can promise is not the resolution time but the time of the next report. "We will report back in an hour, organizing what we know and what we do not" can always be kept, and every time it is kept, trust builds.

Twenty minutes of silence doubles the outage in the customer's imagination. So the first message during an outage must include the cadence. "We will report every 30 minutes until recovery."

How it works

The structure of an incident report has a different order from everyday documents. In everyday documents you write the background first, but an incident report starts with the impact.

영향     : 결제 API 5xx 비율 평시 0.1% → 최대 7% (14:10–15:40)
현재 상태: 완화 조치 적용 완료, 오류율 평시 범위로 복귀 확인
잠정 원인: 14:05 설정 배포에서 DB 커넥션 풀 크기 축소
다음 단계: 원복 완료. 재발 방지로 배포 전 설정 diff 점검 절차를 제안 예정

What the reader needs to know first is "who cannot do what right now" and "is it all right now." Root cause analysis, however curious you are, comes after. With a report that starts from the cause, you must read three paragraphs to learn the current state, and in a document an executive reads, those three paragraphs are not read.

And there are three more rules.

Write it in numbers. Not "many users were affected" but "500 responses went from a usual 33 to 70 in one minute." An adjective is interpreted differently by each reader, and that difference becomes a dispute later.

Do not make a person the subject. Not "because the owner changed the configuration wrongly" but "after the configuration change." If a specific owner is named in a customer report, at the next outage that person hides information. A blameless report is not morality but an investment that buys the speed of the next diagnosis.

Translate the jargon. Not "the Pod has fallen into a restart loop" but "the server is repeatedly turning off and on, so connections are being cut." The same goes for status codes. The number 500 means something only to engineers, and the customer needs "payments failed."

What it looks like in the field

In the same situation, sentences that erode trust and sentences that build it diverge.

When you do not yet know the cause, "I don't think it is a problem on our side" is defensive. "So far we have confirmed the network and authentication are normal, and we are looking at the data layer. We will share an interim result in 30 minutes" describes the same state while building trust.

The difference is three elements. Confirmed facts, what we are doing now, and the time of the next report. If these three are in it, a report holds even in a state where you do not know the cause.

Finally, escalation. It is not a confession of failure. If you decide in advance your own rule of escalating when there is no progress within 30 minutes, the decision to escalate becomes a procedure and not an emotion. In front of the customer, "we have brought in specialist staff" is read not as a signal of incompetence but as a signal of response.

The conversation after an outage ends

When recovery is done, the nature of the conversation changes. During the outage it was "what is happening now," but after it ends it becomes "why did it happen, and will it not happen again?" The document that answers this question must be different from the reports during the outage, and it must come out within days. After a week people's memories diverge, and fact-checking itself becomes hard.

There are four things to include. A chronological record, the cause, why it was found late, and what will be changed. The third is the one most often missing, but in fact the improvement that comes from it is the most valuable. If it could have been fixed in 30 minutes but took two hours to notice, what needs work is not the procedure for fixing but the device for noticing.

When writing improvement items, who does it by when must be written together. An improvement list with no owner and deadline gets written again identically in the next incident report. And if there are ten items, nothing gets done, so it is more honest to keep only the two or three that are truly effective and write down the rest while stating explicitly that they will not be done.

It is worth pointing out the FDE's position here. Carrying out the improvements is usually the customer's part, and we are the ones who build the basis. So instead of "you must do this," "this value comes out like this, and if you set an alert on it, next time you can know within 30 minutes" is received much better. Giving the material to judge and letting the customer decide ends up making more improvements actually get carried out.

Finally, write down what went well too. If the rollback took only 5 minutes, that is thanks to someone having prepared the procedure in advance, and if that fact is not recorded, the reason to make that kind of preparation next time disappears. If an incident report becomes a document that collects only bad things, people come to dislike writing it, and from then on the record itself disappears.

What you will do in the next lab

Using the values found in the previous three courses as material, you produce one page of an incident report with the required sections in place, with numbers in it, naming no person, and with the customer summary written without technical jargon. Grading looks only at whether those conditions are met, not at whether the sentences are good or bad.