The Stabilisation Period and Handover to SM
In one line
Stabilization is not time for fixing the remaining bugs but time for enabling the operations organization to run the system on its own, and if you misunderstand this, the phone keeps ringing even after the contract ends.
Why this is a problem
Because the urgent and the important are in exactly opposite directions. Right after go-live, outages and inquiries pour in so the day is filled with responding alone, and operations documentation and knowledge transfer keep being pushed back with "just until this week is over".
When the stabilization period ends that way, the SM owner takes over a system they know nothing about. The result is a phone call at two in the morning a few months later — the contract has ended but you are the only person to ask. The result of stabilization should be measured not by "how many outages there were" but by "does it now run without us?"
The stabilization period is usually 3–6 months
The contract has an item called the 'stabilization period'. It is the window right after go-live in which the contractor is responsible for incident response and fixing early defects. Many newcomers misunderstand the nature of this period. Stabilization is not 'time for fixing the remaining bugs' but 'time for enabling the operations organization to run the system on its own'.
So the work in the stabilization period has three strands.
- Actual incident and inquiry response (the urgent)
- Operations documentation — operator manual, incident response procedure, batch operations guide (the important)
- Knowledge transfer — training of the SM owner, side-by-side working, handover confirmation (a contractual obligation)
If you cling only to the urgent, the SM owner is left knowing nothing on the day stabilization ends. And 6 months later at 2 a.m. you get a phone call. The contract has ended, but the call comes.
What actually happens in the first week after go-live
| When | Typical issue |
|---|---|
| Morning of go-live day | A login surge — all employees connect at once. Connection pool and session settings are the first hurdle |
| Days 1–2 | Screen error inquiries. Mostly data problems (missing code values, anomalies in migrated data) |
| Day 3 | Abnormal result from the first night batch. Boundary conditions of the data migration |
| Week 1 | First time the month-end/weekly work is performed. Inquiries surge: "Where is this screen?" |
| Month 1 | First month-end close. Reports that statistics and settlement numbers don't match |
A pattern is visible. Most issues right after go-live are not code defects but data and configuration. That is why the data verification scripts and configuration backups made during cutover are used throughout the stabilization period.
The standard flow of incident response
In the field, leaving this order usually makes things worse.
1. 접수·기록 언제, 누가, 무슨 화면에서, 어떤 메시지
2. 영향 범위 파악 전체인가 일부인가. 특정 사용자/특정 데이터만인가
3. 임시 조치 서비스 복구 우선 (재기동, 우회, 기능 임시 차단)
4. 원인 분석 로그·모니터링·최근 변경 이력
5. 항구 조치 코드/데이터/설정 수정과 배포
6. 재발 방지 모니터링 추가, 검증 로직 추가, 문서 갱신
Don't swap the order of steps 3 and 4. An attitude of acting only after fully understanding the cause lengthens the outage. But while doing step 3, always preserve the evidence — if you don't capture a thread dump and logs before the restart, you will never find the cause. If "it worked once we restarted" repeats three times, the fourth time even a restart won't fix it.
The defect log and the 'defect vs requirement' fight
The biggest conflict in the stabilization period is this.
Customer — "This doesn't work, so it's a defect. Please fix it."
Contractor — "That wasn't in the requirements. It's additional development."
What serves as the basis in this fight is the requirements definition document and the RTM. So the document you made in month 1 protects the company in month 8.
Manage the defect log like this.
- Defect ID / date registered / registrant / symptom / reproduction steps / severity (critical, medium, minor) / owner / date of action / action taken / judgment (defect/request/inquiry)
- The judgment column is the key. If you fix everything without judging, the project never ends.
Severity criteria are also agreed in advance. Usually 'critical' means work is halted, 'medium' means a workaround is possible, and 'minor' means an inconvenience. The response time (SLA) differs by severity.
What to hand over when moving to SM (maintenance)
A list goes into the handover confirmation. In practice, at least this much should be there.
- System configuration diagram (servers, ports, accounts, paths — matching reality)
- Deployment procedure and rollback procedure (ones actually followed once)
- Batch list: schedule, predecessor/successor relationships, action on failure
- Interface list: counterpart systems, contact persons, order of contact in an outage
- Account list: DB, WAS, OS, external systems (and who manages the passwords)
- List of known issues and temporary workarounds — you must not hide this. It will come out within 3 weeks anyway
- Monitoring items and thresholds
What does work in the maintenance phase look like
SM is not 'fixing things when they break'. The actual share of work is roughly this.
- 40% change requests (legal amendments, business rule changes, added screens)
- 25% regular tasks (batch monitoring, month/quarter-end close support, backup checks)
- 20% handling inquiries
- 15% incident response
So the ability an SM owner needs is not 'fast development' but 'the ability to judge the scope of impact accurately'. A good SM person is one who knows that a request to add one column affects three integrated systems. The basis for that judgment is in the end the interface definition document and the table definition document. Documents earn their real value after the project ends.
What you see in the field
The first week after go-live follows an almost fixed order.
On the morning of the day, logins are concentrated. All employees connect at the same time, so connection pool and session settings are the first hurdle, and if it gets stuck here, it looks as if the go-live itself has failed, not just the system. On days 1–2, screen error inquiries come in, and a good many of them are not defects but "it's different from the old system". If you accept these as defects, the defect log grows to hundreds of entries in a few days, and the real defects get buried in it.
So the most important document in this period is the classification criteria of the defect log. If you decide case by case whether something is a defect or a new requirement, it becomes a fight each time, but if the criteria exist first, judging becomes routine administration.
The same goes for handover. Training crammed into the last week does not stick. The true indicator of handover is how many times the SM owner has handled real incidents together with you.