RPO, RTO and 3-2-1
In one line
Every choice in backup design comes from two numbers: RPO (how much can we afford to lose?) and RTO (how quickly must we be back?).
Why this was needed
"How often should we run backups?" has no answer. Rephrase it as "How many minutes of data can we afford to lose?" and an answer appears.
- RPO (Recovery Point Objective) — how far back in time you can bring data at the moment of failure. In other words, the acceptable amount of data loss. If the RPO is 1 hour, the backup interval must also be 1 hour or less.
- RTO (Recovery Time Objective) — the time allowed from the failure until the service resumes. This is determined not only by the backup method but by the entire recovery procedure.
The two numbers are directly tied to cost. Getting the RPO down to minutes requires continuous archiving, and getting the RTO down to minutes requires a standby system. So these numbers are set by the business, not by technology. The engineer's job is to take those numbers and turn them into a feasible design.
How it works
Three methods
| Method | What it holds | Backup time | Recovery time | Storage |
|---|---|---|---|---|
| Full | Everything | Long | Short (unpack just one) | Large |
| Incremental | Changes since the last backup | Short | Long (the full backup plus every incremental, in order) | Small |
| Differential | Changes since the last full backup | Medium | Medium (the full backup plus the latest differential) | Medium |
The difference between incremental and differential is easy to confuse, because the reference point differs. Incremental is relative to the "previous backup," and differential is relative to the "last full backup." So a differential grows larger day by day, but you need only two pieces to recover.
A common combination in practice is a weekly full backup plus daily incrementals. The risk of this combination, however, is that if one incremental is corrupted, everything after it is invalid. That is why a generation retention policy and regular verification come as a pair.
The 3-2-1 rule and its modern additions
The classic rule: three copies, on two different kinds of media, with one of them off-site.
Two more are commonly added today.
- One is offline or immutable — a countermeasure against ransomware. If the backup server is reachable from the production network, the backups get encrypted along with everything else.
- 0 errors — verify recovery regularly to confirm that the error count is 0.
Consistency — data that changes during a backup
If an application is writing to a file while you copy it, the backup becomes a mix that matches no point in time. This is especially fatal for databases.
There are three layers of solutions.
- Application-level dump —
pg_dump -Fc,mysqldump --single-transaction. The most reliable and the most portable. - Filesystem snapshot — LVM/Btrfs/ZFS. Freeze a point in time, then back up that frozen state. But it is only a block-level point-in-time freeze, so data the application was holding in memory is not reflected. And when the snapshot volume fills up, the snapshot is invalidated, so size it generously.
- Continuous archiving — keep archiving WAL/binary logs to allow recovery to an arbitrary point in time. Use it when the RPO must be reduced to minutes.
What a recovery drill actually reveals
The part of a backup strategy that goes unverified is always the recovery side. If you actually bring a system back to life once in a while, you will usually run into several of the following.
There is nowhere to recover to. You have the backup, but no disk or server to unpack it onto. The time it takes to obtain those resources during an outage is added directly to the RTO. To measure recovery time for real, you must start from an empty server.
What you need is not in the backup. You got the database, but the uploaded files are missing, or the configuration files, or the certificates and keys. It happens because, when making the list, you thought only of "data" and did not count everything needed to bring the service back up.
The key is inside the backup. If you keep the decryption key for an encrypted backup only in the same place as that backup, it is useless. Conversely, if you keep the key where no one knows about it, recovery is impossible. Keep the key somewhere other than the backup, but where two or more people can reach it.
No one knows the order. The documentation does not say whether the database should come up first, whether the cache should be flushed, or what to do about the messages left in the queue. So the recovery succeeds, but the service does not return to normal.
The backup had silently stopped. If you only send success notifications, a state in which nothing arrives looks normal. Measure the age of the latest backup as a metric and alert on it. "Alert if there is no successful backup within 24 hours" is stronger than "alert on failure."
No one has ever checked that what was received is readable. Many setups count it as a success if the file size matches. At a minimum, unpack it, and if it is a database, actually connect and run a one-line query. A backup whose recovery has never been tested is not a backup but a file.
What you see in the field
It helps to remember the five ways a backup silently breaks.
- It is missing from the targets — a new volume was added but was not put into the backup script.
- You see only the success log and miss the failure — cron is silent by default. You need a mechanism that checks the exit code and reports failures.
- Ransomware gets replicated into the backups too — a setup that keeps only the latest backup is helpless in this situation.
- The restore target environment is different — if the UID/GID mapping, SELinux context, or kernel version differs, the restore works but the service does not start.
- The backup hurts production performance, so the backup window shrinks — then you lengthen the interval, and at some point you can no longer meet the RPO.
The minimum line of defense for item 2 is this.
#!/usr/bin/env bash
set -euo pipefail
trap 'echo "backup FAILED at line $LINENO" >&2; exit 1' ERR
rsync -aHAX --numeric-ids --delete /srv/data/ /backup/prod/data/
echo "backup OK $(date -Is)"
And it is far better to record the success time in a file and monitor "is the latest success within 24 hours?" than to read logs.
What to look for in the next check
The quiz that follows checks the conditions under which you choose RPO and RTO, full, incremental, and differential backups, and off-site copies. Once you pass that bar, the following modules move on to tar full and incremental backups, rsync generation retention, verification scripts, and timed recovery rehearsals.