Capacity Planning and Change Management — Calculate When It Fills, Write Down When to Stop
Measured, Not Promised — RPO, RTO, 3-2-1 and Blameless Postmortems
In one line
RPO and RTO are goals you write in a document, and a recovery drill measures in how many minutes you actually kept those goals. Actual data loss is counted from "the last successful backup," and actual recovery time is counted from the moment the outage started to the moment the service was confirmed. When the drill is over, there should be a blameless post-incident analysis and remediation actions that have an owner and a deadline.
Why you need this
The sentence "backups run every hour, and the RPO is 1 hour" is often false. If the backup job failed twice in a row and nobody knew, the actual loss is three hours. "Recovery takes 30 minutes" also measures only the time the restore command runs, while it takes 15 minutes to notice the outage, 13 minutes to decide on recovery, and 13 minutes to prepare, separately. The gap between the goal and the measurement is visible only by doing a drill.
How it works
The two goals. NIST SP 800-34 Rev. 1 defines the terms of contingency planning as follows. RTO (Recovery Time Objective) is the maximum time a system may be down, and RPO (Recovery Point Objective) is up to which point before the outage the data must be restored — that is, the maximum span of time of data you may lose.
| Goal | How to measure | |
|---|---|---|
| RPO | The maximum time you may lose | Outage time − the time of the last successful backup before the outage |
| RTO | The maximum time you may be down | Service confirmation time − outage start time |
There are two places where measurement often goes wrong. One is counting RPO from "the last backup time" without checking whether that backup failed, and the other is counting RTO as "from the start of the restore to its end." To the users, the service had been down since the moment the outage started.
Break it into phases. If you divide recovery time into the phases detection → decision → preparation → restoration → verification, it shows where to cut. If the restore command is the longest, fix the backup method or bandwidth; if detection is long, fix the alerting; if the decision is long, fix the procedure of who decides what.
The 3-2-1 rule. This is a rule that the US-CERT recommended in 2012 in Data Backup Options — keep three copies of the data (including the original), on two different kinds of media, with one of them at a different location. Three snapshots inside the same disk array have three copies but one medium and one location, so they violate the rule. If the array dies entirely or the building catches fire, all three disappear together.
Verify the restored copy. That a restore is "finished" and that the data is "correct" are different facts. If at backup time you also leave a list (a manifest) containing a hash for each file, then after restoring you can compare against the list and tell apart missing files from files with different contents. sha256sum -c does that job.
Blameless post-incident analysis. The chapter on postmortem culture in the SRE book says a post-incident analysis must be blameless. Once you start looking for who was at fault, people hide facts and the same conditions remain. In the root cause section, you write not "who" but the conditions that left that person no choice but to do it that way — the backup failure did not lead to an alert, nobody was watching the free space of the volume being backed up. You attach an owner and a deadline to each remediation action. An action with no owner is done by nobody.
What it looks like in the field
The fact that most often surfaces in a drill is that the backups had been failing silently. The volume being backed up was full and the incrementals failed, but the alert went only by email and nobody looked at that mailbox — such stories are common. So backup monitoring should not stop at "alert if a job fails" but should be set as "alert if the last success is older than N minutes." A failure alert does not come when the job does not run at all.
The second is that the recovery procedure exists only in people's heads. A drill works best when someone who does not know the procedure tries it using only the runbook. If the reason the preparation phase was long in the drill is "looking for the connection information for the restore server," then putting that information in the runbook is the remediation action. The third is doing the drill once and stopping. Data size and composition keep changing, so a restore that took 45 minutes half a year ago may take two hours now. If you repeat the same drill every quarter and accumulate the per-phase times in a table, whether you can meet the goals shows up as a trend.
What you will do in the next lab
From the copy lists of four sets of data, you find the ones that violate 3-2-1. You measure the actual data loss time and recovery time from the backup job record and the incident record, compare them with the goals, and find the longest phase. You compare the restored files against the manifest to separate what is missing from what has changed, and finally you write a blameless post-incident analysis and remediation actions with owners and deadlines.