TT Lab
Get started
Learn Learning paths Courses

Kubernetes Operations

The Recovery a Snapshot Cannot Do

Continue in TT Lab

Summary

An etcd snapshot is a tool that rolls back the whole cluster to a single point in time, so you cannot use it for "just this namespace" or "to another cluster." Object-level backup fills that place, but you must not put the exported objects back as they are, and in particular, if you leave ownerReferences in, the restored objects quietly disappear.

Why snapshots alone are not enough

The requests that actually come in operations usually look like this. "Please bring back just the shop namespace that I deleted by mistake yesterday afternoon." Or "Please build the staging cluster with the same configuration as production." Neither request can be answered with an etcd snapshot. A snapshot restore rolls the whole cluster back to that point in time, so every other change that happened in between disappears too. You cannot undo a whole day's deployments for one deleted namespace. And a snapshot is tied to the identity of its cluster, so it cannot be moved to another cluster.

It is more accurate to see the two approaches as having different places.

etcd snapshot Object backup
Unit The whole cluster Chosen objects
Point in time One point for everything The time each object was exported
To another cluster Not possible Possible
Partial restore Not possible Possible
Control plane configuration Restored along with it Not covered

How it works

The core of object backup is not exporting but cleaning. kubectl get -o yaml exports everything without distinguishing what the user wrote from what the server filled in. Some of it must not be put back.

What to remove Why
metadata.uid A new one is issued for every object. You cannot put in the old value
metadata.resourceVersion A value for optimistic locking, so on restore it only causes conflicts
metadata.creationTimestamp, generation Decided by the server
metadata.managedFields A history of field ownership, meaningless in another cluster
metadata.ownerReferences In the restored cluster there is no owner with that uid
status The controller observes it and writes it again
Service's spec.clusterIP, nodePort Addresses that cluster allocated

The habit of also removing metadata.namespace is useful. It lets you decide where to restore in the command rather than in the manifest, and that is what lets you build a procedure of first standing it up in a verification namespace with the same backup.

The most dangerous one is ownerReferences. Kubernetes manages dependencies through this reference, and the garbage collector deletes objects whose owner has disappeared. But the reference includes not just the name but also the uid. If the restored cluster cannot find an object with that uid, the garbage collector decides the owner is gone and deletes the restored object. No error, no notification, and in most cases not even an event. This is where the incident comes from in which you back up a whole ReplicaSet and it quietly disappears a few seconds after the restore.

Restores also have an order. The namespace must exist before you can put objects in it, and the CRD must be registered first before it will accept custom resources. If you get the order wrong, the API server rejects it, saying it does not know the kind at all — the manifest is not wrong; that kind just does not exist yet.

Finally, a PVC is not data. If you backed up a PersistentVolumeClaim, what you restored is only a request form saying "give me this much volume," and the files that were in it are nowhere. Data backup is a different problem, of volume snapshots or application-level backups. If you blur this distinction, you can pass a recovery drill and still lose data in a real incident.

What it looks like in the field

The most common incident is a restore that disappears a few seconds later. It is because of the ownerReferences above, but the symptom is not "the restore failed" but "the restore succeeded and it is gone," so finding the cause takes a long time.

The second is not writing down the restore order. Normally the CRD already exists so there is no problem, and you meet it for the first time on the day you restore everything into a new cluster. That day there is no time.

The third is verifying by a person's eyes. If you only count whether objects exist, a restore in which the replicas dropped from 2 to 1 also looks like a success. That is why the last step of a restore procedure must always be a machine comparing values.

Limits of this lab environment

This cluster has no PersistentVolume and no CSI driver, so it cannot show the difference between a PVC and data in practice. That part is covered by the explanation above. Conversely, the garbage collector really runs, so if you restore with a stale owner reference left in, you can see that object actually disappear — when I measured it in the lab, it was gone in a few seconds. Also, this lab never uses etcdctl. It is meant not to overlap with the etcd lab in the same course, and object backup really is a task that does not touch etcd.

What to do in the next lab

You create one CRD and one set of a namespace (Deployment, Service, ConfigMap, and a custom resource), and export them raw to see what is in them. Then you build a script that strips the fields that must be removed, delete the namespace entirely, and bring it back into a different namespace. Next you confirm that an object left with a stale owner reference actually disappears, read the error when a custom resource goes in before its CRD, and write down the restore order. Finally you build a verification script that compares the restore list with the actual cluster.

References: