TT Lab
Get started
Learn Learning paths Courses

Kubernetes Operations

Taking an etcd Snapshot and Verifying It

Continue in TT Lab

Goal

You check the state of a live etcd, take a snapshot, and get hands-on practice with the procedure for mechanically verifying that the snapshot is actually a usable thing.

Why it matters

Making a backup that you can trust is harder than making a backup. In reality there are many systems in which a CronJob that runs every day leaves a 0-byte file every day and no alarm ever rings. That is why this lab separates the step of creating the file from the step of judging that file. snapshot status returns three values, revision, totalKey, and hash, and they are each an answer to "what point in time is it," "how much does it hold," and "is it intact." If you make a script read these three and print OK, the success of a backup is expressed by an exit code rather than by a person's eyes. One more thing: engrave into your mind, together with the member count calculation, that quorum prevents failures but does not prevent accidental deletion. Whether there are three members or five, a wrong deletion is faithfully replicated to all members.

Steps

  1. Save the output of etcdctl --endpoints=127.0.0.1:2379 endpoint health to /root/etcd/backup/out/health.txt. The file must show healthy and the endpoint port 2379, and must not contain the word unhealthy.
  2. Save the output of etcdctl member list to /root/etcd/backup/out/members.txt (it must include the hexadecimal member ID and the 2379/2380 addresses). Then, in /root/etcd/backup/out/quorum.txt, write two lines, members=<멤버 수> (number of members) and quorum=<과반 수> (majority count). The etcd in this lab is a single member, so it becomes members=1 and quorum=1.
  3. Save a snapshot to /root/etcd/backup/snap-01.db. The file size must be at least 20000 bytes.
  4. Run snapshot status in JSON format and save the result to /root/etcd/backup/out/snap-01.json. revision must be greater than 0, totalKey greater than 20, and hash not 0.
  5. Count the keys under the /registry prefix and put only the number in /root/etcd/backup/out/key-count.txt (it must be 20 or more). Then save a top list of key counts per resource type to /root/etcd/backup/out/top-prefixes.txt. This file must contain the name of at least one of pods, configmaps, secrets, leases, events, and namespaces.
  6. Create the namespace etcd-lab and in it create the ConfigMap etcd-lab-marker with data.owner=platform. After that, take a second snapshot to /root/etcd/backup/snap-02.db and save the verification result to /root/etcd/backup/out/snap-02.json. The revision of the second JSON must be greater than that of the first.
  7. In /root/etcd/backup/cronjob.yaml, write the spec for a periodic backup. It is kind: CronJob, metadata.namespace: kube-system, and spec.schedule: "0 */6 * * *". The container's command is in list form and must contain snapshot save, the four flags --endpoints, --cacert, --cert, and --key, and a delete command that uses find and -mtime +7. The Pod spec must have a volume whose hostPath is /etc/kubernetes/pki/etcd, a node selector key that contains control-plane (nodeSelector), and tolerations. Write all volumes in hostPath form.
  8. Make /root/etcd/backup/verify.sh an executable script. If the snapshot passed as an argument is healthy, it must print OK <revision> <keys> (space-separated, two numbers) on the first line, and otherwise print a line that does not start with OK. With this script, save the result for snap-01.db to /root/etcd/backup/out/verify-01.txt and the result for snap-02.db to /root/etcd/backup/out/verify-02.txt. Finally, deliberately create a corrupted file, verify it, and leave the result in /root/etcd/backup/out/verify-bad.txt.

Reference

Check etcd health before the backup

Save the output of etcdctl --endpoints=127.0.0.1:2379 endpoint health to /root/etcd/backup/out/health.txt. The file must show healthy and the endpoint port 2379, and must not contain the word unhealthy.

If you back up a sick etcd, you get a sick snapshot. Use the subcommand of the endpoint family of etcdctl that checks health, and leave the output in a file. This lab has no client TLS, so no certificate flags are needed.

The member list and computing quorum

Save the output of etcdctl member list to /root/etcd/backup/out/members.txt (it must include the hexadecimal member ID and the 2379/2380 addresses). Then, in /root/etcd/backup/out/quorum.txt, write two lines, members=<멤버 수> (number of members) and quorum=<과반 수> (majority count). The etcd in this lab is a single member, so it becomes members=1 and quorum=1.

Look at the actual members with member list, and quorum is the majority. Remember that the majority of N members is the quotient of N divided by 2, plus 1.

Take the first snapshot

Save a snapshot to /root/etcd/backup/snap-01.db. The file size must be at least 20000 bytes.

Put the file path to save to after snapshot save. Check that the resulting file size is tens of KB or more — a 0-byte backup is a common incident.

Verify the snapshot as JSON

Run snapshot status in JSON format and save the result to /root/etcd/backup/out/snap-01.json. revision must be greater than 0, totalKey greater than 20, and hash not 0.

Having taken one is different from having a usable one. snapshot status with -w json prints the hash, revision, and totalKey as one line of JSON. The warning message goes to standard error, so it does not get mixed into the redirection.

Count what is in etcd and how much

Count the keys under the /registry prefix and put only the number in /root/etcd/backup/out/key-count.txt (it must be 20 or more). Then save a top list of key counts per resource type to /root/etcd/backup/out/top-prefixes.txt. This file must contain the name of at least one of pods, configmaps, secrets, leases, events, and namespaces.

All Kubernetes objects are under the /registry prefix. Give get the prefix query option and the keys-only option, and the resource type is the third piece of the key path.

Make a change and take a second snapshot

Create the namespace etcd-lab and in it create the ConfigMap etcd-lab-marker with data.owner=platform. After that, take a second snapshot to /root/etcd/backup/snap-02.db and save the verification result to /root/etcd/backup/out/snap-02.json. The revision of the second JSON must be greater than that of the first.

Creating one object raises the revision. If you compare the revisions of the two snapshots, the meaning of a backup's point in time becomes tangible. The order, taking it after creating, is the key.

Write the periodic backup CronJob spec

In /root/etcd/backup/cronjob.yaml, write the spec for a periodic backup. It is kind: CronJob, metadata.namespace: kube-system, and spec.schedule: "0 */6 * * *". The container's command is in list form and must contain snapshot save, the four flags --endpoints, --cacert, --cert, and --key, and a delete command that uses find and -mtime +7. The Pod spec must have a volume whose hostPath is /etc/kubernetes/pki/etcd, a node selector key that contains control-plane (nodeSelector), and tolerations. Write all volumes in hostPath form.

etcd exists only on the control plane, and there is a taint on it. Mount the certificates from a host path, and put a retention policy that deletes old files into the command too.

Build a snapshot verification script

Make /root/etcd/backup/verify.sh an executable script. If the snapshot passed as an argument is healthy, it must print OK <revision> <keys> (space-separated, two numbers) on the first line, and otherwise print a line that does not start with OK. With this script, save the result for snap-01.db to /root/etcd/backup/out/verify-01.txt and the result for snap-02.db to /root/etcd/backup/out/verify-02.txt. Finally, deliberately create a corrupted file, verify it, and leave the result in /root/etcd/backup/out/verify-bad.txt.

For a healthy file, print OK in the specified format, and for a corrupted file you must not print OK. If a verifier always passes, it is verifying nothing.