TT Lab
Get started
Learn Learning paths Courses

Backup — Bring Deleted Data Back

There Are Two Kinds of Backup

Continue in TT Lab

One-line summary

A logical backup is a photograph; a physical backup plus WAL is a video. A photograph only takes you back to "last night's state," while a video can take you back to "30 seconds before the incident."

Why you need it — decide two numbers first

Designing backups is not about picking a tool; it is about answering two numbers.

If you take a pg_dump once a day, your RPO is 24 hours. In the worst case, a full day of data is gone. The first question is whether you can live with that; if you cannot, you need WAL archiving.

Do not forget RTO either. If restoring a 500 GB dump takes 6 hours, the service is down for 6 hours even though you have a backup.

Logical backup — pg_dump

pg_dump -Fc -d labdb -f labdb.dump      # 커스텀 포맷 (압축·병렬 복원 가능)
pg_restore -d newdb labdb.dump

But it is only a snapshot of the moment it was taken. There is no way to get back to a point between last night's dump and this afternoon's incident. And on a large database, both dumping and restoring can take hours.

A dump is for "moving house" and "partial recovery," not the main tool for disaster recovery.

Physical backup + WAL — PITR

PostgreSQL writes every change to the WAL (Write-Ahead Log) first. The WAL comes before the data files. That makes the following equation hold.

어느 시점의 데이터 파일 복사본  +  그 뒤의 WAL 전부  =  그 뒤 아무 시점이나

This is PITR (Point-In-Time Recovery). It takes three pieces.

  1. archive_mode = on — copy each finished WAL segment to a safe place
  2. Base backup — the whole data directory, taken with pg_basebackup
  3. Recovery settings — feed the WAL back in with restore_command and stop at recovery_target_time

Rules for archive_command

archive_command = 'test ! -f /archive/%f && cp %p /archive/%f'

When it fails, PostgreSQL retries and does not delete the WAL. If the archive destination fills up, pg_wal keeps growing until the disk is full. If you do not watch pg_stat_archiver.failed_count, it grows quietly until it blows up.

Recovery procedure

# 1. 베이스 백업을 복사한다 (원본은 건드리지 않는다)
cp -r /backup/base /var/lib/postgresql/restore

# 2. 어디서 멈출지 알려 준다
cat > restore/postgresql.auto.conf <<EOF
restore_command = 'cp /archive/%f %p'
recovery_target_time = '2026-08-21 16:20:41+00'
recovery_target_action = 'promote'
port = 5433
EOF

# 3. 복구 모드로 시작하라는 표시
touch restore/recovery.signal

# 4. 띄운다
pg_ctl -D restore start

If recovery.signal exists, PostgreSQL starts in recovery mode. It feeds the WAL back in, and when it reaches the target time it does whatever recovery_target_action says.

What you must observe when recovering

Never recover on top of the original. Start the recovery in a different directory and on a different port, check it, and only then move it into place. That way you keep room to back out if you picked the wrong target time.

The timeline forks. When you promote after recovery, the timeline number goes up by 1 (00000002...). From that point on, the original and the recovered copy are different histories. If you do not know this, you will later get lost over "the WAL doesn't match."

Common mistakes

Keeping the backup on the same disk. If the disk dies, both die. A backup must be on a different machine, preferably in a different location.

Not practicing restores. This is the most common and the most expensive. A green backup job and a backup you can actually recover from are two different things. Unless you run recovery rehearsals regularly, you cannot say you have a backup.

Not watching archive_command failures. As described above, the disk fills up.

Keeping only logical backups. That means your RPO is one day, and usually nobody ever agreed to that.

In practice, you use a tool

Instead of writing your own scripts, you use pgBackRest or Barman. They include incremental backups, parallel compression, retention policies, S3 upload, and above all backup verification. This lab is meant to show what those tools do inside.

In one sentence:

A backup you have never recovered from is not a backup.