etcd — No Snapshot, No Cluster
Summary
Everything you see with kubectl get is a value read from etcd. If you lose etcd, you lose the cluster.
Why this matters
Growing the control plane to several machines is reassuring. There is a record of the author growing a homelab from 3 nodes to 7, making the control plane 3 machines, and setting the etcd members to 3. With 3 members, the majority is 2, so the cluster stays alive even if one machine dies. But the last line of that write-up has this note. Quorum is for failures, not for mistakes.
This sentence is the starting point of this module. Replication protects you from an incident in which a node dies. But replication cannot protect you from the incident of mistakenly typing kubectl delete ns payments. A deletion is a normal write, and the three etcd members faithfully replicate that deletion three times over. With 100 members, the result is the same. The only means of turning back time is a snapshot from before that point.
The same write-up has another case. With all 3 control plane machines healthy, when the first node died, both kubectl and the kubelets lost their connections. This was because controlPlaneEndpoint was hard-coded to the first node's physical IP rather than a VIP. etcd quorum and API availability were separate problems. The feeling that "we made it redundant so we're fine" is often wrong like this. Only someone who has done the recovery procedure by hand once knows what is actually protected.
How it works
An etcd backup is not copying a file but taking a whole, consistent point in time.
etcdctl snapshot save /backup/snap.db \
--endpoints=https://127.0.0.1:2379 \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key
Production etcd accepts requests only over TLS with client authentication on. That is why three certificates always come along. You do not need to memorize the paths. In a kubeadm cluster etcd runs as a static Pod, so the answer is written as is in the execution arguments of kubectl describe pod etcd-<노드> -n kube-system (where the placeholder is the node name). The etcd in this lab environment listens in plaintext at 127.0.0.1:2379 without client TLS, so the three flags are omitted. Instead, in the step of writing a periodic backup CronJob, you will write the production form as it is.
After you take one, always verify it.
etcdutl snapshot status /backup/snap.db -w json
# {"hash":3106878859,"revision":12450,"totalKey":1287,"totalSize":5779456}
All three numbers carry meaning. revision is the point in time the snapshot holds, totalKey is the scale of the objects it holds, and hash is file integrity. Backup systems that stack up 0-byte files every day without complaint are more common than you would think. An unverified backup is not a backup but a hope for one.
A restore is the opposite direction, but there is one decisive difference. snapshot restore is not a command that rolls data back but a command that creates a new etcd cluster. That is why it takes identity flags such as --name, --initial-cluster, --initial-advertise-peer-urls, and --initial-cluster-token. And --data-dir must be a new, empty path. If you give it the existing /var/lib/etcd as is, it fails or, worse, touches live data.
The order of a restore is also fixed.
| Order | What you do | Why |
|---|---|---|
| 1 | Verify the snapshot | If you build a cluster from an unusable file, it dies twice |
| 2 | Stop kube-apiserver | Writes must not continue during the restore |
| 3 | snapshot restore | Creates a new data directory |
| 4 | Replace the hostPath in the etcd manifest with the new path | If you miss this, it comes back up with the old data |
| 5 | Start the apiserver and check with kubectl get |
Whether it came back to life is seen through the API |
What it looks like in the field
First, the backup interval is the maximum loss. If you back up every 6 hours, in the worst case 6 hours' worth of changes disappear. You set the recovery point objective (RPO) and then set the interval to match it, not the other way around.
Second, the mistake of keeping the backup file on the same disk. When a node's disk dies, the snapshot dies with it. In the author's homelab too, a separate path out to the NAS was built. A retention policy is also needed — without a line like find /backup -mtime +7 -delete, the disk fills up a few months later.
Third, certificates and time after a restore. A snapshot holds only objects. The nodes' kubelet certificates, tokens, and the actual containers created in the meantime are a world outside the snapshot. It is normal for the cluster to be messy right after a restore, and after a while the controllers converge it to the declared state.
What to do in the next lab
In the first lab you check the health and members of a live etcd, take a snapshot twice and watch the revision increase, and build a periodic backup CronJob and a snapshot verification script. In the second lab you actually delete an object, restore the snapshot into a new data directory, and start a second etcd process yourself with that data to see with your own eyes the deleted value come back to life.