The Three Stages of Verification
In one line
Backup verification has three stages: does it exist → is it readable → does it recover. The first two can be automated, and the last has to be done by a person.
Why this was needed
Backups run every day. The log says success. Then, on the day you actually need to recover, the file will not open. This situation is not rare.
How it works
Stage 1 — does it exist
This is the most basic check, and surprisingly many problems are caught here.
ls -lh /backup/prod/ | tail -5
find /backup/prod -type f -mtime -1 | wc -l
Confirm that the file size is not 0 and that the timestamp is recent. Having monitoring watch "is the latest success within 24 hours?" is far better than reading logs.
Stage 2 — is it readable
Check the compression integrity and the archive listing.
gzip -t /backup/etc-2026-08-20.tar.gz && echo 'gzip OK'
tar -tzf /backup/etc-2026-08-20.tar.gz > /dev/null && echo 'archive OK'
sha256sum -c /backup/checksums.sha256
The three each check something different. gzip -t checks the CRC of the compressed stream, tar -t checks the archive structure, and sha256sum -c checks the hash of the whole file. There are cases where the compression is fine but the tar structure is broken, and vice versa.
Stage 3 — does it actually recover
Only this is real verification. Restore to another host or a temporary environment, bring up the service, and confirm that the data is from the point in time you expected.
Rehearsal checklist.
- Are the credentials for accessing the backup storage kept outside the production server?
- Is the restore procedure document somewhere other than the restorer's laptop?
- Is the encryption key needed for the restore kept separately?
- Does the actual time the rehearsal took come in within the RTO?
- After the restore, does the application start and pass the consistency check?
Storing the encryption key is especially often forgotten. If you encrypt the backups and keep the key only on the server being backed up, then when that server disappears, the backups become meaningless binary files.
Once per quarter is the realistic minimum for the rehearsal interval. And a rehearsal should be performed by someone other than the person who made the backup, using only the documentation, so that the gaps in the procedure are exposed.
The shape of a verification script
#!/usr/bin/env bash
set -euo pipefail
ARCHIVE="${1:?usage: verify.sh <archive>}"
[ -s "$ARCHIVE" ] || { echo "빈 파일이거나 없음: $ARCHIVE" >&2; exit 1; }
gzip -t "$ARCHIVE" || { echo "압축 스트림 손상" >&2; exit 2; }
tar -tzf "$ARCHIVE" > /dev/null || { echo "아카이브 구조 손상" >&2; exit 3; }
echo "OK $ARCHIVE"
The knack is to give a different exit code for each stage. Monitoring can then tell immediately at which stage it failed.
What you see in the field
Bit rot. An archive stored for a long time is silently damaged. The disk itself reports as healthy, yet a few bytes have flipped. Regular checksum re-verification catches this. One reason to use ZFS or Btrfs as backup storage is their built-in checksums.
Silent failures on tape and object storage. There are cases where a write reports success but was not actually stored. You need a procedure that reads back and checks right after the write.
Why 3-2-1 is still valid today
It is an old rule, but it holds unchanged in the cloud.
사본 3개 · 매체 2종 · 오프사이트 1개
Two more are commonly added today.
- One is offline or immutable — having ransomware encrypt the backups as well is the standard tactic. There must be one copy that even administrators cannot delete, such as Object Lock in COMPLIANCE mode.
- 0 unverified copies — a backup you have not checked does not count as a copy.
The second is the point of this module. Rather than boasting about how many backups you have, what matters is whether there is at least one backup that you have restored.
Automating integrity verification
Verification that depends on a person remembering to do it will stop someday. Run three layers automatically.
# 1) 파일이 온전한가 — 매일
sha256sum -c backup-2026-09-06.sha256
gzip -t backup-2026-09-06.sql.gz # 압축 무결성만 확인(빠르다)
# 2) 읽을 수 있는가 — 주마다
pg_restore --list backup.dump > /dev/null # 목록만 뽑아 본다
# 3) 복원되는가 — 분기마다
pg_restore -d verify_db backup.dump && psql verify_db -c "select count(*) from orders"
Layer 1 takes seconds, layer 2 takes minutes, and layer 3 takes hours. Setting different intervals according to cost is the practical answer. If you do none of the three, all that remains is the belief that "we have backups."
Where to record verification results
If verification fails and no one knows, it is the same as not having done it. Provide three things.
- Metric — export the last success time as a gauge. Alert if
time() - last_success > 2일(the expression uses a Korean unit for 2 days). The time elapsed since the last success is a more accurate signal than the failure count. - Record — when, what, and how you verified, and what the result was. Audits require this.
- Alert — keep failures of the backup itself and failures of verification as separate alerts. The causes and responses differ.
# 백업이 돌지 않는다 (심각)
time() - backup_last_success_timestamp > 86400 * 2
# 백업은 도는데 검증이 실패한다 (더 심각 — 있다고 믿었던 것이 없다)
backup_verify_success == 0
The second is more serious. Noticing that a backup is not running is easy, but a state in which it is running yet unusable goes unnoticed by anyone until the moment it is needed.
What you will do in the next lab
You create a checksum manifest, introduce damage on purpose and detect it, and write a verification script with a distinct exit code per stage. The grader runs that script against both a healthy archive and a corrupted archive.