TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Backing Up etcd and Planning the Restore

Continue in TT Lab

Goal

You actually take and verify an etcd snapshot, confirm through a resource comparison what point in time a snapshot captures, and leave the recovery procedure as a document.

Why it matters

An etcd backup comes up almost every time on the CKA hands-on exam. But memorizing the command is not enough. What this lab emphasizes is the fact that a snapshot captures a point in time. A resource created after the snapshot is not in that file, and it disappears when you restore. Without this sense, you get the incident "why is the thing I created yesterday missing?" after a restore.

There is also a reason to have you write the recovery plan in words. A restore is not one command but an order. You bring the control plane down, restore into a new data-dir, match the member configuration and peer URLs, and bring it back up. If you do not write this order in advance, you end up making it in the middle of an outage, and the judgment then is usually wrong.

The etcd in this environment runs at 127.0.0.1:2379. On a real kubeadm cluster, the three options --cacert, --cert, and --key would be added here.

Steps

  1. Create /root/cka-etcd/env.sh and export ETCDCTL_API=3 and ETCDCTL_ENDPOINTS=127.0.0.1:2379 so that they are passed to child processes.
  2. Save the etcd endpoint status (or health) output to /root/cka-etcd/status.txt. The file must show the endpoint address.
  3. Create the namespace cka-etcd and the ConfigMap pre-backup. The data is stage=before. Then save the list of ConfigMaps in cka-etcd to /root/cka-etcd/before.txt.
  4. Save an etcd snapshot to /root/cka-etcd/snap.db.
  5. Save the status (metadata) output of that snapshot to /root/cka-etcd/snap-status.txt.
  6. After taking the snapshot, create the ConfigMap post-backup (data stage=after) in cka-etcd, and save the ConfigMap list to /root/cka-etcd/after.txt. before.txt must not contain post-backup, and after.txt must contain both.
  7. In /root/cka-etcd/restore-plan.md, write the restore procedure in at least five lines and at least 200 bytes. It must contain all of the words snapshot restore, --data-dir, --initial-cluster, the peer port 2380, and apiserver, and it must also say that the control plane is briefly stopped during the restore.

Reference

Set up the etcdctl environment variables

Create /root/cka-etcd/env.sh and export ETCDCTL_API=3 and ETCDCTL_ENDPOINTS=127.0.0.1:2379 so that they are passed to child processes.

etcdctl has a split between the v2 and v3 APIs. Environment variables must be passed to child processes, so simply assigning them is not enough.

Check the endpoint status

Save the etcd endpoint status (or health) output to /root/cka-etcd/status.txt. The file must show the endpoint address.

Use both endpoint status and endpoint health. The output sometimes goes to standard error, so be careful with redirection.

Create a backup reference point

Create the namespace cka-etcd and the ConfigMap pre-backup. The data is stage=before. Then save the list of ConfigMaps in cka-etcd to /root/cka-etcd/before.txt.

To compare later which point in time the snapshot captures, you need to leave the current state in a file. This file must lack something that does not exist yet, so that the later steps hold.

Save a snapshot

Save an etcd snapshot to /root/cka-etcd/snap.db.

snapshot save takes a file path as an argument. It fails if the directory does not exist, so create it first.

Check the snapshot metadata

Save the status (metadata) output of that snapshot to /root/cka-etcd/snap-status.txt.

status shows the hash, revision, key count, and size. Recent etcd moved this subcommand to a separate tool, so if it is missing, try that one.

Compare with resources created after the backup

After taking the snapshot, create the ConfigMap post-backup (data stage=after) in cka-etcd, and save the ConfigMap list to /root/cka-etcd/after.txt. before.txt must not contain post-backup, and after.txt must contain both.

A resource created after the snapshot is not in that snapshot. This step proves that fact through the difference between the two files.

Putting it together: leave the recovery procedure as a document

In /root/cka-etcd/restore-plan.md, write the restore procedure in at least five lines and at least 200 bytes. It must contain all of the words snapshot restore, --data-dir, --initial-cluster, the peer port 2380, and apiserver, and it must also say that the control plane is briefly stopped during the restore.

A restore is not one command but an order. Write what you stop first, where you restore to, and which values you must match for it to come back up.