A Backup You Have Never Restored
Summary
A backup you have never restored from is not a backup but a file you believe is a backup. The value of a backup is proven only by recovery.
Why this is a problem
Backups do not fail noisily. The script produces errors but nobody looks at the exit code, so 0-byte files pile up every day; the target table list is stale, so only newly created tables are missing; or the backup file is on the same disk as the original and disappears together with it the day the disk dies.
In all of these cases, the daily inspection report says "backup OK." So the problem is first discovered on the very day the backup is needed.
What often catches people is everything other than data. A backup missing sequences, permissions, and triggers will not bring the application up even after recovery. If you run a recovery rehearsal even once, you learn this then; if you do not, you learn it on the day of the outage.
The value of a backup is proven only by recovery
In an operations review meeting, you hear the report "backups are running normally every day." There should be only one follow-up question.
"When was the last time you restored from that backup?"
If the answer is "never," that is not a backup but a file you believe is a backup. Things like these actually happen.
- The backup script is producing errors but nobody checks the exit code, so 0-byte files are created every day
- The backup works, but the target table list is stale, so newly created tables are missing
- The backup file is on the same disk, so it disappears together when the disk fails
- Recovery takes 3 hours, and you learn that on the day of the outage
- Recovery works, but sequences, permissions, and triggers are missing so the application does not come up
The last item is especially common. Recovery does not work just because you have the data.
The three axes of a backup
| Axis | Question | Metric |
|---|---|---|
| What | Only data? Schema, permissions, and sequences too? | The backup target list |
| How often | How much can we afford to lose | RPO (recovery point objective) |
| How fast | How long can we afford to be stopped | RTO (recovery time objective) |
RPO and RTO are not technology but values the business decides. You have to ask the business owner, "Is it acceptable to lose one hour of data?" The backup frequency and method are decided by that answer.
- RPO 24 hours → one full backup a day is enough
- RPO 1 hour → full backup + incremental/log backup
- RPO close to 0 → replication + log point-in-time recovery
Most SI projects do not have this conversation. So all there is is the technical fact that "backups run every dawn," and nobody knows whether that satisfies the business requirement.
Backup during cutover — time is the option
A backup for a cutover (launch) job has a different nature. Because it must be usable as a rollback means.
In a job that starts at 2 a.m. and has to make a decision by 4 a.m., a backup that takes 3 hours to recover is not an option. So these two lines must be in every cutover plan.
백업 소요 시간 : 실측 __분
복구 소요 시간 : 실측 __분 ← 이 값이 롤백 판단 시한을 결정한다
If recovery takes long, prepare other means.
- Snapshots (storage/virtualization level) — usually much faster
- A separate backup of only the tables to be changed — if you do not need to revert everything
- Design schema changes to be reversible (expand-migrate-contract)
What is often missing from backup targets
If you back up only the database and stop there, recovery does not work.
- DB accounts and permissions — if the recovered DB has no accounts, the application cannot connect
- Current values of sequences/auto-increment — if they go back, PK collisions occur
- Views, triggers, procedures, and functions
- Application configuration files — the very ones emphasized in the cutover module as well
- Batch schedule definitions
- Certificates and keys
So you need to manage the backup target list as a document and have a procedure to update the list when new objects are added. If the automatic backup captures the "whole schema," a lot is solved, but you only know that if you have ever checked.
The recovery rehearsal procedure
You run the rehearsal like this. The trick is to do it somewhere other than the production server, with only the backup.
1. 백업 파일을 별도 서버/디렉터리로 가져온다 (운영에서 직접 복구하지 않는다)
2. 빈 상태에서 복구를 수행한다 — 시작·종료 시각을 기록
3. 검증
- 테이블 수, 주요 테이블 행 수
- 시퀀스 현재값
- 계정·권한·인덱스·제약
- 애플리케이션 기동과 로그인 (가능하면)
4. 결과를 문서로 — 소요 시간, 발견된 누락, 조치
5. 발견된 누락을 백업 대상 목록에 반영
The last line of item 3 is the real value of the rehearsal. Discovering "recovery worked but the application does not come up" in the rehearsal and not on the day of the outage.
Automating backup verification
It is hard for a person to do a rehearsal every time. At a minimum, automate this much.
# 매일 백업 직후 자동 검증
1. 백업 파일이 생성됐는가 (존재 + 크기 > 최소 기준)
2. 백업 명령의 종료코드가 0인가 ← 놀랍게도 이걸 안 보는 곳이 많다
3. 파일이 열리는가 (압축이면 무결성 검사)
4. 예상 객체가 포함돼 있는가 (테이블 목록 grep)
5. 주 1회: 실제 복구 후 행 수 비교
I want to emphasize item 2. If the backup script fails and cron swallows the error, nobody knows. You must check the exit code and send an alert on failure. "We have never received a backup failure notification" does not mean "backups always succeed" but may mean "there is no notification at all."
Retention and deletion
Backups pile up without limit. You need a policy.
일 백업: 14일 보관
주 백업: 8주 보관
월 백업: 12개월 보관
The retention count is the range you can go back. With 14 daily backups you can go back two weeks. If it is a line of business that might get a request such as "please restore the data from 3 months ago," you must keep it for that period. This is also an item to agree with the business.
And automate deletion too, but with safeguards. There are real cases where a bug in an "delete old ones" script deleted everything. Delete only by date condition, set a minimum retention count as a floor, and leave the list in the log before deleting.
And, keep backups elsewhere
A backup on the same server and the same disk dies together with that server. Keep a copy on at least a different storage, and if possible in a different physical location.
In an air-gapped environment, an external cloud is not an option, so you use in-house dedicated backup storage or tape/external media. In that case, media transfer procedures and encryption come with it. This is also an item to settle before cutover.
What it looks like in the field
There is a moment when planning a cutover when the backup becomes a time calculation. If the backup takes 3 hours to recover, the rollback decision deadline must be 3 hours before the end of the work. If you do not do this arithmetic in advance, at 4 a.m. you realize on the spot that "if we roll back now we cannot make the start of the morning business."
So what to write in the backup item is not "backup complete" but the file path and the recovery duration. The rule of having the actual file name written in the cutover result report came from the same reason — so that you do not spend time at the critical moment searching for which file it is.