Undoing a Deletion Incident With a Snapshot
Goal
You actually delete an object and then bring it back from a snapshot. You start a second etcd yourself with the restored data directory and see with your own eyes that "it really came back."
Why it matters
There is a big difference between knowing the recovery procedure only from documents and having done it once. In particular, the fact that snapshot restore is not a command that rolls data back but a command that creates a new etcd cluster stays in your body only if you do it by hand. That is why it requires identity flags such as --name, --initial-cluster, --initial-advertise-peer-urls, and --initial-cluster-token, and demands a new, empty path for --data-dir. The moment you give it the existing /var/lib/etcd as is, you end up touching live data. Another important thing is the order. You verify the snapshot first, stop the apiserver to cut off writes, restore, change the manifest's hostPath to the new path, and finally check with kubectl. If you reverse this order, you end up building a cluster from an unusable snapshot, losing writes that arrived during the restore, or coming back up with the old data. In this lab, without stopping the live cluster, you bring up the restored copy on a separate port and reach the same conclusion.
Steps
- Create the namespace
etcd-laband the ConfigMapetcd-lab-marker(data.owner=platform) in it. Then record the pre-incident state as JSON in/root/etcd/restore/out/pre.json. There are three keys:marker_owner(valueplatform),registry_keys(the number of keys under/registry, 20 or more), andrevision(the current etcd revision, greater than 0). - Take the snapshot to use for the restore to
/root/etcd/restore/before.db(at least 20000 bytes), and save the verification result as JSON to/root/etcd/restore/out/before.json.revisionmust be greater than 0. - Delete the ConfigMap
etcd-lab-markerto reproduce the incident. After the deletion, leave the output of querying that object in/root/etcd/restore/out/gone.txt. The file must contain a "does not exist" response such asNotFoundornot found, and this ConfigMap must not reappear in the live cluster. - Restore
before.dbto/root/etcd/restore/data. After the restore,/root/etcd/restore/data/member/snap/dband/root/etcd/restore/data/member/walmust exist. Save the output of the restore command (including standard error) to/root/etcd/restore/out/restore.txt. - This time attach all the identity flags and restore once more to
/root/etcd/restore/data2. All five flags--data-dir,--name,--initial-cluster,--initial-advertise-peer-urls, and--initial-cluster-tokenmust be included, with the namerestored, the peer addresshttp://127.0.0.1:12380, and the tokenetcd-restore-lab. Leave the entire command you actually ran, as is, in/root/etcd/restore/out/restore-cmd.txt(--data-dirmust not be given/var/lib/etcd). - Start a second etcd process in the background that uses
/root/etcd/restore/data2as its data directory. The client address ishttp://127.0.0.1:12379, the peer address ishttp://127.0.0.1:12380, the name isrestored, the initial member isrestored=http://127.0.0.1:12380, and the token isetcd-restore-lab. Do not terminate this process until grading is finished. Save the command you used to start it to/root/etcd/restore/out/second-etcd.txt(the file must show12379and--data-dir). - From the restored copy, read the key
/registry/configmaps/etcd-lab/etcd-lab-marker, check that the value containsplatform, and save that output to/root/etcd/restore/out/recovered.txt. The query target is127.0.0.1:12379, and this object must still be absent from the live cluster (2379). - In
/root/etcd/restore/runbook.md, write the recovery runbook in at least 400 bytes. Put each step on a different line and keep the order below. (1) Verify the snapshot (snapshot status), (2) stopkube-apiserver, (3)snapshot restore, (4) replace the etcd manifest'shostPath/data-dirwith the new path, (5) check withkubectl get. At the end, be sure to also write how to roll back if the recovery fails.
Reference
- Snapshot verification and restore are standard with
etcdutl. First check withcommand -v etcdutl; if it exists useetcdutl snapshot status/etcdutl snapshot restore, and if not, substitute the same subcommands ofetcdctl(a warning is attached but it works). - The etcd server executable is separate. Look first with
command -v etcd, and if it is not there, find it withfind "$HOME/.kwok" -name etcd -type f(the binary that kwokctl downloaded is under it). For a background start, the formnohup <etcd 경로> <플래그들> > /root/etcd/restore/out/second-etcd.log 2>&1 &(where the placeholders are the etcd path and the flags) is convenient. - If you make the ports of the second etcd overlap with the live etcd (2379/2380), startup fails. Be sure to use 12379/12380, and keep the process alive until steps 6 and 7 are graded.
- Both the restore log and the
kubectlNotFound message go to standard error. You have to capture them with명령 > 파일 2>&1or명령 2>&1 | tee 파일(the placeholders stand for the command and the file) for them to remain in the file. - Values stored in etcd are protobuf, so they look garbled as they are. If you strip only the null bytes with
| tr -d '\0', the strings inside read as they are. - Common mistake 1: writing
etcdctl snapshot restore --data-dir ...on one line in the runbook of step 8. Thensnapshot restoreanddata-dirend up on the same line and get caught by the order check. Split it across several lines or put the hostPath replacement as a separate item. For the same reason, do not mentiondata-dir,kubectl get, andkube-apiserverearly in the document — the order is judged by the first position where each word appears. - Common mistake 2: recreating the live cluster's ConfigMap in step 3. The conclusion of this lab is "it comes back only in the restored copy."
- The lab Pod comes up fresh for each lab, so the
etcd-labnamespace and the marker created in the earlier lab do not remain. Create them yourself in step 1 and start from there. This is exactly why operational procedures must be left in runbooks and manifests rather than in memory.
Record the pre-incident state as numbers
Create the namespace etcd-lab and the ConfigMap etcd-lab-marker (data.owner=platform) in it. Then record the pre-incident state as JSON in /root/etcd/restore/out/pre.json. There are three keys: marker_owner (value platform), registry_keys (the number of keys under /registry, 20 or more), and revision (the current etcd revision, greater than 0).
A recovery starts with knowing 'what was there.' Bundle three things into one JSON: the marker value, the number of /registry keys, and the current revision. The revision is in the header of endpoint status.
Secure and verify the snapshot for the restore
Take the snapshot to use for the restore to /root/etcd/restore/before.db (at least 20000 bytes), and save the verification result as JSON to /root/etcd/restore/out/before.json. revision must be greater than 0.
You verify a recovery snapshot as soon as you take it. If you learn that the file is corrupt only after the incident, it is already too late.
Reproduce the deletion incident and leave a trace
Delete the ConfigMap etcd-lab-marker to reproduce the incident. After the deletion, leave the output of querying that object in /root/etcd/restore/out/gone.txt. The file must contain a "does not exist" response such as NotFound or not found, and this ConfigMap must not reappear in the live cluster.
Delete the object and then query it. The message for querying a nonexistent resource goes to standard error, so you must include standard error in the redirection for it to remain in the file.
Restore the snapshot into a new data directory
Restore before.db to /root/etcd/restore/data. After the restore, /root/etcd/restore/data/member/snap/db and /root/etcd/restore/data/member/wal must exist. Save the output of the restore command (including standard error) to /root/etcd/restore/out/restore.txt.
A restore is not overwriting an existing path but creating a new one. The restore log also goes to standard error. Check whether the member structure was created inside the resulting directory.
Restore again with the cluster identity flags attached
This time attach all the identity flags and restore once more to /root/etcd/restore/data2. All five flags --data-dir, --name, --initial-cluster, --initial-advertise-peer-urls, and --initial-cluster-token must be included, with the name restored, the peer address http://127.0.0.1:12380, and the token etcd-restore-lab. Leave the entire command you actually ran, as is, in /root/etcd/restore/out/restore-cmd.txt (--data-dir must not be given /var/lib/etcd).
A restore is actually a command that 'creates a new cluster.' If you specify the four things, name, initial member list, peer address, and token, you can then start etcd with that data.
Start a second etcd with the restored copy
Start a second etcd process in the background that uses /root/etcd/restore/data2 as its data directory. The client address is http://127.0.0.1:12379, the peer address is http://127.0.0.1:12380, the name is restored, the initial member is restored=http://127.0.0.1:12380, and the token is etcd-restore-lab. Do not terminate this process until grading is finished. Save the command you used to start it to /root/etcd/restore/out/second-etcd.txt (the file must show 12379 and --data-dir).
You can verify only by attaching the restored data directory to a real process. Keep the client and peer ports from overlapping with the live etcd, and keep the process alive until grading is finished.
Check that the deleted object is alive in the restored copy
From the restored copy, read the key /registry/configmaps/etcd-lab/etcd-lab-marker, check that the value contains platform, and save that output to /root/etcd/restore/out/recovered.txt. The query target is 127.0.0.1:12379, and this object must still be absent from the live cluster (2379).
The etcd key of a Kubernetes object is /registry/<종류>/<네임스페이스>/<이름> (type, namespace, name). The value is protobuf, so it is hard for a person to read, but the strings show as they are.
Write the recovery runbook
In /root/etcd/restore/runbook.md, write the recovery runbook in at least 400 bytes. Put each step on a different line and keep the order below. (1) Verify the snapshot (snapshot status), (2) stop kube-apiserver, (3) snapshot restore, (4) replace the etcd manifest's hostPath/data-dir with the new path, (5) check with kubectl get. At the end, be sure to also write how to roll back if the recovery fails.
This is a document you will read at 3 a.m. The order of the steps matters, and you must also write where to go back to if it fails.