Capacity Planning and Change Management — Calculate When It Fills, Write Down When to Stop
Measure RPO and RTO from a Recovery Drill and Write the Postmortem
Goal
Find the 3-2-1 violations in the backup copy list, measure the actual data loss time (RPO) and recovery time (RTO) from the backup job record and the incident record and compare them with the goals, compare the restored copy against the manifest, and then write a blameless post-incident analysis and remediation actions with owners and deadlines.
Why it matters
"RPO 1 hour, RTO 90 minutes" is just a goal in a document. If the backups had been failing silently, the actual loss is several times larger, and if you measure only the restore command time, the time spent on detection and decision drops out. A recovery drill is the work of revealing that gap as numbers, and a post-incident analysis is the document that turns those numbers into the next improvement. Once you start writing down who was at fault, the facts go into hiding and the conditions remain as they are.
The material is in /opt/lab/capacity/dr/, and the grader computes the expected values from the original material.
Steps
- Copy all the contents of
/opt/lab/capacity/dr/to/root/dr/. - Write the data that violates the 3-2-1 rule (3 copies, 2 kinds of media, 2 locations) to
/root/dr/01-321.txtasviolations=. - Write
last_good=,loss_min=, andmeets_rpo=to/root/dr/02-rpo.txt. - Write
rto_min=,meets_rto=, andlongest_phase=to/root/dr/03-rto.txt. - Compare the restored copy against
MANIFEST.sha256and writemissing=andmismatch=to/root/dr/04-verify.txt. - Write five sections to
/root/dr/postmortem.md: summary, impact, timeline, root cause, and remediation actions (six times in the timeline, the two numbers in the impact, and a root cause without any person's name). - Write three or more remediation actions, each with
담당:and기한: YYYY-MM-DD(owner and deadline labels attached), and cover the detection, backup, and 3-2-1 gaps all together.
Notes
- Time calculation:
datetime.strptime(s, '%Y-%m-%d %H:%M'), and the difference of two times.total_seconds() // 60. - Restore verification:
cd restored && sha256sum -c MANIFEST.sha256. - Common mistake 1: counting a failed backup as last_good — it is a point you cannot restore to.
- Common mistake 2: counting RTO from the start of the restore — the users had been down since the moment of the outage.
Copy the material
Copy all the contents of /opt/lab/capacity/dr/ (including the restored/ directory) to /root/dr/.
To move a whole directory, use cp -r. Start by reading incident.log and backup_jobs.log.
Check the 3-2-1 rule
For each data set in backups.json, look at its copy list (including the original), and write the data that violates even one of 3 or more copies, 2 or more kinds of media, and 2 or more locations (site) to /root/dr/01-321.txt as violations=<이름들, 쉼표로> (the names, separated by commas).
For each data set, count len(copies), the number of distinct media, and the number of distinct sites. Three snapshots at the same location violate the rule even though there are three copies.
Actual data loss time
Using the failure time in incident.log and backup_jobs.log, write three lines to /root/dr/02-rpo.txt: last_good=<장애 이전에 OK 로 끝난 마지막 백업, YYYY-MM-DD HH:MM> (the last backup that ended OK before the outage), loss_min=<장애 시각 − last_good, 분> (the outage time minus last_good, in minutes), and meets_rpo=<targets.env 의 RPO_MIN 이하면 yes, 아니면 no> (yes if at or below RPO_MIN in targets.env, otherwise no).
It is not the last backup but the last successful backup. A job that ended as FAILED cannot be a point to restore to.
Actual recovery time and the longest phase
Using incident.log, write three lines to /root/dr/03-rto.txt: rto_min=<failure 부터 verified 까지, 분> (from failure to verified, in minutes), meets_rto=<RTO_MIN 이하면 yes, 아니면 no> (yes if at or below RTO_MIN, otherwise no), and longest_phase=<detect·decide·prepare·restore·verify 가운데 가장 긴 구간> (the longest of detect, decide, prepare, restore, and verify). The phases are, in order, failure→detected, detected→declared, declared→restore_start, restore_start→restore_end, and restore_end→verified.
Count not only the time the restore command ran but from the moment the outage started. To the users, the service had been down since then.
Verify the restored copy
Compare the files in /root/dr/restored/ against MANIFEST.sha256 and write two lines to /root/dr/04-verify.txt: missing=<목록에는 있는데 복원본에 없는 파일, 쉼표로> (files that are in the list but not in the restored copy, separated by commas) and mismatch=<있지만 해시가 다른 파일, 쉼표로> (files that are there but have a different hash, separated by commas).
If you run sha256sum -c MANIFEST.sha256 inside the restored directory, each file gets OK, FAILED, or missing. Even if it ends with an error, read the output to the end.
A blameless post-incident analysis
Write /root/dr/postmortem.md. It must have five sections, ## 요약, ## 영향, ## 타임라인, ## 근본 원인, and ## 개선 조치 (in order: summary, impact, timeline, root cause, and remediation actions). In the timeline, transfer all six times from incident.log, and in the impact, write the actual loss minutes and recovery minutes you measured in steps 3 and 4 as numbers. In the root cause, do not write the account name of a person in charge; write the conditions that caused the outage and the conditions that delayed recovery.
Write not who made the mistake but the conditions that left that person no choice but to do it that way (there was no alert, nobody was watching the free space). You may fill in the remediation actions section in step 7, but the section heading must be there now.
Remediation actions with owners and deadlines
In the ## 개선 조치 section (the remediation actions section) of postmortem.md, write three or more items that start with - . Each item must have 담당: <누구> and 기한: YYYY-MM-DD (owner and deadline labels), and there must be at least one action each for the three gaps this incident revealed — detection (alerting), backup (the failure was silent), and 3-2-1 (the violation you found in step 2).
An action with no owner is done by nobody. "Alert if the last successful backup is older than N minutes" catches even the case where the job does not run at all, better than "alert on backup failure."