TT Lab
Get started
Learn Learning paths Courses

Backup and Restore

Designing the Restore Procedure

Continue in TT Lab

In one line

Recovery is all about order and verification. And until you actually walk through that order and time it, the RTO is not a number but a wish.

Why this was needed

Even when backups are made well, recovery usually fails because of problems with the procedure.

How it works

A safe restore order

# 1. 임시 경로로 먼저 푼다 — 절대 바로 덮어쓰지 않는다
mkdir -p /restore/check
tar -xzf /backup/full-2026-08-20.tar.gz -C /restore/check

# 2. 내용을 확인한다
find /restore/check -type f | wc -l
diff -r /restore/check/etc/nginx /etc/nginx | head -40

# 3. 필요한 것만 옮긴다
rsync -aHAX --numeric-ids /restore/check/etc/nginx/ /etc/nginx/

Going through a temporary path is the key. Extracting directly with -C / is an irreversible operation.

The order of an incremental restore

You must keep the order: full → incremental 0 → incremental 1 → ... Incrementals made with tar's --listed-incremental also carry deletion information, so if you break the order, files that should exist disappear.

tar --listed-incremental=/dev/null -xzf full.tar.gz -C /restore/
tar --listed-incremental=/dev/null -xzf inc0.tar.gz -C /restore/
tar --listed-incremental=/dev/null -xzf inc1.tar.gz -C /restore/

Using --listed-incremental=/dev/null when restoring is the convention. It means you will not update the snapshot file.

Selective restore

You can pull out specific files without extracting everything.

tar -tzf backup.tar.gz | grep nginx.conf
tar -xzf backup.tar.gz -C /restore/ src/conf/nginx.conf
tar -xzf backup.tar.gz -C /restore/ --strip-components=1 src/conf/

--strip-components=N removes the first N path elements when extracting. Use it when the archive structure and the restore location differ.

Timing it

To confirm the RTO, you have to actually measure it.

START=$(date +%s)
# ... 복구 절차 전체 ...
END=$(date +%s)
echo "복구 소요: $((END - START))초"

What must be included is not only the archive extraction time. The time to fetch the files from backup storage, the verification time, the service start-up time, and the consistency check time must all be included to give the real RTO. The number one reason RTO is missed in practice is not slow extraction but the time to bring the backup over the network.

Verification after the restore

Having a backup versus being able to recover

A log saying the backup succeeded only means the file was created. You have to confirm three things before you can say "we can recover."

  1. Is it readable? — Is the compression intact, and is the encryption key still alive?
  2. Is it complete? — Does it contain everything that is needed (schema, data, sequences, extensions)?
  3. Does it finish in time? — When you actually restore, it takes much longer than expected.

The third is the one that often collapses. If restoring 500GB takes 6 hours and the RTO is 1 hour, that backup cannot be used for disaster recovery. To meet the RTO you need a replica or a snapshot — a logical backup is the last line of defense, not the first measure.

Method Restore time Recovery point Used for
Logical dump (pg_dump) Slow (hours) Time of the dump Migration, partial restore
Physical backup + WAL Moderate (minutes to hours) Any point in time Disaster recovery
Storage snapshot Fast (minutes) Time of the snapshot Quick rollback
Promoting a read replica Fastest (seconds) Nearly real time High availability

What you actually measure in a rehearsal

Do not stop at "the recovery worked"; leave numbers behind.

## 복구 리허설 2026-09-06

- 대상: labhub-prod DB (실 용량 142GB)
- 백업 시각: 2026-09-05 12:30 KST
- 복원 목표 시점: 2026-09-05 18:00 KST (PITR)

| 단계 | 걸린 시간 |
|---|---|
| 백업 내려받기 | 21분 |
| 기본 백업 복원 | 47분 |
| WAL 재생 (5.5시간분) | 33분 |
| 서비스 기동·검증 | 9분 |
| **합계 (RTO 실측)** | **1시간 50분** |

- 검증: 행 수 대조 3개 표 일치, 최근 주문 10건 육안 확인, 애플리케이션 기동 성공
- 발견: WAL 아카이브에 12분 구멍(2026-09-05 14:02~14:14) — 아카이빙 재시도 설정 필요
- RPO 실측: 12분 (목표 5분 미달 — 개선 필요)

A rehearsal with no findings is about as good as no rehearsal. The first time you do one, almost always something turns up.

When to do it

And do the rehearsal in a place isolated from production. The moment you restore onto the production DB, it is no longer a rehearsal but an incident.

What you see in the field

"The restore worked but the service does not start." The candidate causes are well known. UID/GID mapping, SELinux context, differences in filesystem features (no ACL support), and differences in kernel/library versions. --numeric-ids and --selinux are the safeguards for the first two.

The restore document exists only on the restorer's laptop. The situation in which that laptop is unavailable may be exactly the situation in which recovery is needed.

What you will do in the next lab

You carry out a recovery rehearsal with the backup set you built earlier. You restore in a safe order, compare with the original, do selective restore and incremental-order restore, and measure the time to determine whether the RTO is met.