Backup — Bring Deleted Data Back
There Are Two Kinds of Backup
One-line summary
A logical backup is a photograph; a physical backup plus WAL is a video. A photograph only takes you back to "last night's state," while a video can take you back to "30 seconds before the incident."
Why you need it — decide two numbers first
Designing backups is not about picking a tool; it is about answering two numbers.
- RPO (recovery point objective) — how much data you can afford to lose
- RTO (recovery time objective) — how quickly you must be serving again
If you take a pg_dump once a day, your RPO is 24 hours. In the worst case, a full day of data is gone. The first question is whether you can live with that; if you cannot, you need WAL archiving.
Do not forget RTO either. If restoring a 500 GB dump takes 6 hours, the service is down for 6 hours even though you have a backup.
Logical backup — pg_dump
pg_dump -Fc -d labdb -f labdb.dump # 커스텀 포맷 (압축·병렬 복원 가능)
pg_restore -d newdb labdb.dump
- It crosses versions and architectures — a dump taken on 16 can be loaded into 17
- You can restore just part of it — a single table, for example
- It is human-readable (
-Fp)
But it is only a snapshot of the moment it was taken. There is no way to get back to a point between last night's dump and this afternoon's incident. And on a large database, both dumping and restoring can take hours.
A dump is for "moving house" and "partial recovery," not the main tool for disaster recovery.
Physical backup + WAL — PITR
PostgreSQL writes every change to the WAL (Write-Ahead Log) first. The WAL comes before the data files. That makes the following equation hold.
어느 시점의 데이터 파일 복사본 + 그 뒤의 WAL 전부 = 그 뒤 아무 시점이나
This is PITR (Point-In-Time Recovery). It takes three pieces.
archive_mode = on— copy each finished WAL segment to a safe place- Base backup — the whole data directory, taken with
pg_basebackup - Recovery settings — feed the WAL back in with
restore_commandand stop atrecovery_target_time
Rules for archive_command
archive_command = 'test ! -f /archive/%f && cp %p /archive/%f'
%p— the source path,%f— the file name- It must return 0 on success and a non-zero value on failure. If it lies and claims success, PostgreSQL deletes that WAL, and that stretch becomes unrecoverable forever
test ! -fprevents overwriting. Overwriting a file of the same name with different content silently breaks recovery
When it fails, PostgreSQL retries and does not delete the WAL. If the archive destination fills up, pg_wal keeps growing until the disk is full. If you do not watch pg_stat_archiver.failed_count, it grows quietly until it blows up.
Recovery procedure
# 1. 베이스 백업을 복사한다 (원본은 건드리지 않는다)
cp -r /backup/base /var/lib/postgresql/restore
# 2. 어디서 멈출지 알려 준다
cat > restore/postgresql.auto.conf <<EOF
restore_command = 'cp /archive/%f %p'
recovery_target_time = '2026-08-21 16:20:41+00'
recovery_target_action = 'promote'
port = 5433
EOF
# 3. 복구 모드로 시작하라는 표시
touch restore/recovery.signal
# 4. 띄운다
pg_ctl -D restore start
If recovery.signal exists, PostgreSQL starts in recovery mode. It feeds the WAL back in, and when it reaches the target time it does whatever recovery_target_action says.
promote— promote so the server can be used (the default)pause— stop and give you time to check. This is the safer choice if you want to decide after checking
What you must observe when recovering
Never recover on top of the original. Start the recovery in a different directory and on a different port, check it, and only then move it into place. That way you keep room to back out if you picked the wrong target time.
The timeline forks. When you promote after recovery, the timeline number goes up by 1 (00000002...). From that point on, the original and the recovered copy are different histories. If you do not know this, you will later get lost over "the WAL doesn't match."
Common mistakes
Keeping the backup on the same disk. If the disk dies, both die. A backup must be on a different machine, preferably in a different location.
Not practicing restores. This is the most common and the most expensive. A green backup job and a backup you can actually recover from are two different things. Unless you run recovery rehearsals regularly, you cannot say you have a backup.
Not watching archive_command failures. As described above, the disk fills up.
Keeping only logical backups. That means your RPO is one day, and usually nobody ever agreed to that.
In practice, you use a tool
Instead of writing your own scripts, you use pgBackRest or Barman. They include incremental backups, parallel compression, retention policies, S3 upload, and above all backup verification. This lab is meant to show what those tools do inside.
In one sentence:
A backup you have never recovered from is not a backup.