Taking an etcd Snapshot and Verifying It
Goal
You check the state of a live etcd, take a snapshot, and get hands-on practice with the procedure for mechanically verifying that the snapshot is actually a usable thing.
Why it matters
Making a backup that you can trust is harder than making a backup. In reality there are many systems in which a CronJob that runs every day leaves a 0-byte file every day and no alarm ever rings. That is why this lab separates the step of creating the file from the step of judging that file. snapshot status returns three values, revision, totalKey, and hash, and they are each an answer to "what point in time is it," "how much does it hold," and "is it intact." If you make a script read these three and print OK, the success of a backup is expressed by an exit code rather than by a person's eyes. One more thing: engrave into your mind, together with the member count calculation, that quorum prevents failures but does not prevent accidental deletion. Whether there are three members or five, a wrong deletion is faithfully replicated to all members.
Steps
- Save the output of
etcdctl --endpoints=127.0.0.1:2379 endpoint healthto/root/etcd/backup/out/health.txt. The file must showhealthyand the endpoint port2379, and must not contain the wordunhealthy. - Save the output of
etcdctl member listto/root/etcd/backup/out/members.txt(it must include the hexadecimal member ID and the2379/2380addresses). Then, in/root/etcd/backup/out/quorum.txt, write two lines,members=<멤버 수>(number of members) andquorum=<과반 수>(majority count). The etcd in this lab is a single member, so it becomesmembers=1andquorum=1. - Save a snapshot to
/root/etcd/backup/snap-01.db. The file size must be at least 20000 bytes. - Run
snapshot statusin JSON format and save the result to/root/etcd/backup/out/snap-01.json.revisionmust be greater than 0,totalKeygreater than 20, andhashnot 0. - Count the keys under the
/registryprefix and put only the number in/root/etcd/backup/out/key-count.txt(it must be 20 or more). Then save a top list of key counts per resource type to/root/etcd/backup/out/top-prefixes.txt. This file must contain the name of at least one ofpods,configmaps,secrets,leases,events, andnamespaces. - Create the namespace
etcd-laband in it create the ConfigMapetcd-lab-markerwithdata.owner=platform. After that, take a second snapshot to/root/etcd/backup/snap-02.dband save the verification result to/root/etcd/backup/out/snap-02.json. Therevisionof the second JSON must be greater than that of the first. - In
/root/etcd/backup/cronjob.yaml, write the spec for a periodic backup. It iskind: CronJob,metadata.namespace: kube-system, andspec.schedule: "0 */6 * * *". The container'scommandis in list form and must containsnapshot save, the four flags--endpoints,--cacert,--cert, and--key, and a delete command that usesfindand-mtime +7. The Pod spec must have a volume whosehostPathis/etc/kubernetes/pki/etcd, a node selector key that containscontrol-plane(nodeSelector), andtolerations. Write all volumes inhostPathform. - Make
/root/etcd/backup/verify.shan executable script. If the snapshot passed as an argument is healthy, it must printOK <revision> <keys>(space-separated, two numbers) on the first line, and otherwise print a line that does not start withOK. With this script, save the result forsnap-01.dbto/root/etcd/backup/out/verify-01.txtand the result forsnap-02.dbto/root/etcd/backup/out/verify-02.txt. Finally, deliberately create a corrupted file, verify it, and leave the result in/root/etcd/backup/out/verify-bad.txt.
Reference
- The etcd in this lab has no client TLS, so it connects without
--cacert/--cert/--key. In production the three flags are required, and the paths are written as they are in the execution arguments ofkubectl describe pod etcd-<노드> -n kube-system(where the placeholder is the node name). You practice that form in step 7. snapshot statushas moved toetcdutlthese days.etcdutl snapshot status <파일> -w json(where the placeholder is the file) is the recommended form, and usingetcdctlalso works with a warning. The warning goes to standard error, so it does not get mixed into a> 파일redirection (the placeholder is a file).- When counting keys, the
--keys-onlyoutput has blank lines mixed in. Counting only lines that start with/registryis accurate. The resource type is the second piece of/registry/<종류>/...(where the placeholder is the resource type), so extract it withcut -d/ -f3and then dosort | uniq -c | sort -rn. - For a corrupted file, making one short garbage file is enough. Do not overwrite a real snapshot.
- Common mistake 1: taking the second snapshot before creating the ConfigMap in step 6. Then the revision does not increase.
- Common mistake 2: the script in step 8 printing a line starting with
OKeven on the failure path. A verifier like that is the same as not verifying. - The lab Pod comes up fresh for each lab, so the cluster state created in the earlier lab does not remain. Create the objects you need yourself within this lab. This is exactly why operational procedures must be left in runbooks and manifests rather than in memory.
Check etcd health before the backup
Save the output of etcdctl --endpoints=127.0.0.1:2379 endpoint health to /root/etcd/backup/out/health.txt. The file must show healthy and the endpoint port 2379, and must not contain the word unhealthy.
If you back up a sick etcd, you get a sick snapshot. Use the subcommand of the endpoint family of etcdctl that checks health, and leave the output in a file. This lab has no client TLS, so no certificate flags are needed.
The member list and computing quorum
Save the output of etcdctl member list to /root/etcd/backup/out/members.txt (it must include the hexadecimal member ID and the 2379/2380 addresses). Then, in /root/etcd/backup/out/quorum.txt, write two lines, members=<멤버 수> (number of members) and quorum=<과반 수> (majority count). The etcd in this lab is a single member, so it becomes members=1 and quorum=1.
Look at the actual members with member list, and quorum is the majority. Remember that the majority of N members is the quotient of N divided by 2, plus 1.
Take the first snapshot
Save a snapshot to /root/etcd/backup/snap-01.db. The file size must be at least 20000 bytes.
Put the file path to save to after snapshot save. Check that the resulting file size is tens of KB or more — a 0-byte backup is a common incident.
Verify the snapshot as JSON
Run snapshot status in JSON format and save the result to /root/etcd/backup/out/snap-01.json. revision must be greater than 0, totalKey greater than 20, and hash not 0.
Having taken one is different from having a usable one. snapshot status with -w json prints the hash, revision, and totalKey as one line of JSON. The warning message goes to standard error, so it does not get mixed into the redirection.
Count what is in etcd and how much
Count the keys under the /registry prefix and put only the number in /root/etcd/backup/out/key-count.txt (it must be 20 or more). Then save a top list of key counts per resource type to /root/etcd/backup/out/top-prefixes.txt. This file must contain the name of at least one of pods, configmaps, secrets, leases, events, and namespaces.
All Kubernetes objects are under the /registry prefix. Give get the prefix query option and the keys-only option, and the resource type is the third piece of the key path.
Make a change and take a second snapshot
Create the namespace etcd-lab and in it create the ConfigMap etcd-lab-marker with data.owner=platform. After that, take a second snapshot to /root/etcd/backup/snap-02.db and save the verification result to /root/etcd/backup/out/snap-02.json. The revision of the second JSON must be greater than that of the first.
Creating one object raises the revision. If you compare the revisions of the two snapshots, the meaning of a backup's point in time becomes tangible. The order, taking it after creating, is the key.
Write the periodic backup CronJob spec
In /root/etcd/backup/cronjob.yaml, write the spec for a periodic backup. It is kind: CronJob, metadata.namespace: kube-system, and spec.schedule: "0 */6 * * *". The container's command is in list form and must contain snapshot save, the four flags --endpoints, --cacert, --cert, and --key, and a delete command that uses find and -mtime +7. The Pod spec must have a volume whose hostPath is /etc/kubernetes/pki/etcd, a node selector key that contains control-plane (nodeSelector), and tolerations. Write all volumes in hostPath form.
etcd exists only on the control plane, and there is a taint on it. Mount the certificates from a host path, and put a retention policy that deletes old files into the command too.
Build a snapshot verification script
Make /root/etcd/backup/verify.sh an executable script. If the snapshot passed as an argument is healthy, it must print OK <revision> <keys> (space-separated, two numbers) on the first line, and otherwise print a line that does not start with OK. With this script, save the result for snap-01.db to /root/etcd/backup/out/verify-01.txt and the result for snap-02.db to /root/etcd/backup/out/verify-02.txt. Finally, deliberately create a corrupted file, verify it, and leave the result in /root/etcd/backup/out/verify-bad.txt.
For a healthy file, print OK in the specified format, and for a corrupted file you must not print OK. If a verifier always passes, it is verifying nothing.