Cutover — Go Out in a Way You Can Come Back From
In one line
Cutover is not deployment but work in which several teams move in a set order, and its success depends on having decided the rollback in advance as criteria expressed in numbers.
Why decide the criteria in advance
"If there's a problem, let's roll back" is not a plan. At three in the morning when people are tired, the definition of a 'problem' differs from person to person, and trying to judge it on the spot always ends the same way: "let's watch a little longer" repeats until you pass the time when you could still have gone back.
So you write the thresholds, the observation window and the decision deadline in the plan. If a time is fixed in it, like "if the smoke test has not passed by 04:00, roll back", the person deciding at that time does not need to summon courage. They only need to carry out what has already been decided.
Cutover is not 'deployment'
For a developer, deployment means uploading files and restarting the service. In SI, cutover (移行) is much broader than that.
[사전] 이행 계획 승인 · 백업 · 공지 · 서비스 중단 안내 · 배치 정지 · 연동 상대 통보
[본작업] DB 스키마 반영 → 데이터 이관 → 애플리케이션 배포 → 설정 반영 → 기동
[검증] 스모크 테스트 → 핵심 업무 시나리오 → 연동 시스템 상호 확인 → 배치 재개
[종료] 이행 결과 보고 → 상황실 운영 → 롤백 판단 시한 경과
Several teams move at once in a set order. So a cutover has a checklist and a timeline, and each item carries an owner, an estimated duration and an 'expected result'. Without an expected result, nobody can judge whether that item succeeded.
Rollback is a 'criterion', not a plan
"If there's a problem, let's roll back" is not a plan. At 3 a.m., when people are tired, the definition of a 'problem' differs from person to person. So you decide criteria in numbers in advance.
| Metric | Threshold | Observation window | Action |
|---|---|---|---|
| Order API error rate | Above 5% | 10 minutes | Roll back |
| Login response time p95 | Above 3 seconds | 15 minutes | Observe, then reassess |
| Batch delay | Above 60 minutes | Once | Roll back |
| Data consistency mismatch | Even a single one | Immediately | Roll back |
And you set a rollback decision deadline. "If verification has not passed by 4 a.m., we roll back, no matter what." Without this deadline, "let's watch a little longer" repeats even at 6 a.m., and in the end you open for business at the start of the workday with a half-broken system.
Never start without a backup
Item 1 on the cutover checklist is always the backup. And for a backup you must confirm not that it 'was made' but that 'it can be restored'. A backup you have never restored from is not a backup.
At the very least, take care of these three.
- DB backup (and the measured recovery time — a recovery that takes 3 hours is not a rollback option in a night cutover)
- Application artifacts (the previous WAR/JAR plus configuration files)
- Configuration files (Tomcat
server.xml,context.xml, nginx conf, firewall policy, batch schedules)
Incidents where the configuration file backup is forgotten are common. You rolled the application back by tag, but
nobody knows who changed the connector settings in server.xml or the DB connection pool size, or when.
The smoke test — it must finish within 5 minutes
The smoke test right after cutover is not about 'whether all features work' but about 'whether anything is fatally broken', checked within 5 minutes.
- Does the service open its port and respond (health check)?
- Can you log in?
- Is the DB connection alive (one query)?
- Does a call to the integrated counterpart system work (one key interface)?
- Is the error log flooding with new exceptions?
If you make this an automated script, the verdict does not waver even when your hands shake at dawn. A smoke test in which a person clicks through screens gets less accurate the more tired that person is.
Things that often blow up in integration testing
The defects that come out of pre-cutover integration testing follow patterns.
- Environment differences — things that work only in the development environment. The top cause is configuration files and firewalls, and second is the state of the DB data (code values that exist in development but not in production).
- Integration timing — the counterpart system's batch runs at 03:00 and ours runs at 02:50. The documents for both said only "early-morning batch".
- Character set and length — 3 bytes per Korean character (UTF-8) vs 2 bytes (EUC-KR). Trying to put 10 Korean characters into a column of length 20 gets truncated under UTF-8.
- Permissions — The production DB account has narrower permissions than the development one. You find out that
CREATE TEMP TABLEdoesn't work on the day of cutover.
If you rehearse these four once in a verification environment configured identically to production before cutover, most are filtered out. A cutover that skips the rehearsal becomes an improvised performance at dawn.
The cutover result report
When the night work ends, you write a result report. The format differs by company, but it must include:
- Actual start/end times and the difference from plan
- The list of work performed and the result of each
- Issues that occurred and actions taken (including those you could not resolve)
- The location of backup files and the rollback procedure (referred to throughout the stabilization period)
- Items to confirm on the next working day
Writing this document diligently halves the failure analysis in the stabilization period. It is the only document that can answer "what did we change at go-live?"
What you see in the field
The typical cutover failure is not a technical problem but an ordering problem.
- You start data migration without stopping the batches, so the records that arrive during migration are lost.
- You don't notify the integration counterpart, so the counterpart system sends messages during our stop window and they all pile up as failures.
- The backup was taken after the schema change, so the moment you try to go back, there is nothing to go back to.
And in the verification phase, a long smoke test becomes the problem. A smoke test that doesn't finish within 5 minutes eats into the decision deadline and leaves no time to decide whether to roll back. The smoke test only checks "whether the core business runs", and the rest is checked in the war room.
This is also why the cutover result report should contain the actual backup file name. A report that just says "backup complete" cannot tell you which file it is when you actually need it.