TT Lab
Get started
Learn Learning paths Courses

Kubernetes Operations

Please Restore Just the Namespace We Deleted Yesterday

Continue in TT Lab

Goal

You export one set of a namespace with mixed kinds object by object, strip with a script the fields that have meaning only in that cluster, and then bring it back into a different namespace. You actually see a stale owner reference quietly delete the restored copy, check the error when you get the restore order wrong, and make a machine do the restore verification.

Why it matters

An etcd snapshot is a tool that rolls back the whole cluster to a single point in time. So you cannot use it for a request to bring back just one namespace deleted yesterday, and you cannot move it to another cluster either. What you need in such cases is object-level backup, and it has a difficulty that snapshots do not. You must not put exported objects back as they are. uid and resourceVersion have meaning only in that cluster, status is a value the controller will write again, and the uid in ownerReferences does not exist in the restored cluster — this last one is especially nasty. The garbage collector sees it as an object with no owner and deletes it with no error and no event.

Steps

  1. In /root/ops-backup/crd.yaml, write the CustomResourceDefinition widgets.ops.example.com — group ops.example.com, kind Widget (plural widgets), namespaced scope, version v1 (served and storage both true), and the schema has an integer spec.size. Then, in /root/ops-backup/scene.yaml, write four objects of the namespace ops-shop in one file — the Deployment catalog (2 replicas, label app: catalog, image nginx:1.27.3), a ClusterIP Service catalog with the same name (selector app: catalog, port 80), the ConfigMap catalog-config (currency: KRW, page-size: "20"), and the Widget w1 (spec.size 3). Create the namespace and apply everything.
  2. Export each of the four objects as is, without processing under /root/ops-backup/raw/ — deploy.yaml (Deployment catalog), svc.yaml (Service catalog), cm.yaml (ConfigMap catalog-config), widget.yaml (Widget w1). Save the output of kubectl get <종류> <이름> -o yaml (kind and name) without touching it.
  3. Create /root/ops-backup/clean.sh — it reads one YAML file given as an argument and prints the cleaned result to standard output only. What must be removed is, from metadata, uid, resourceVersion, creationTimestamp, generation, managedFields, ownerReferences, and namespace; the top-level status; and the Service's spec.clusterIP, spec.clusterIPs, and spec.ports[].nodePort. With this script, clean the four files in raw/ and save them under the same names in /root/ops-backup/clean/.
  4. Delete the namespace ops-shop entirely, and after confirming it is gone, record it on one line in /root/ops-backup/gone.txt — ops-shop TAB NotFound. Before deleting, first check that all four files in /root/ops-backup/raw/ exist.
  5. Create the namespace ops-shop-restored and restore the four files in /root/ops-backup/clean/ into that namespace (the cleaned manifests have no namespace, so you specify it in the command). Wait until the Deployment's 2 Pods are Running, then save it as four lines to /root/ops-backup/restored.tsv — Deployment TAB catalog, Service TAB catalog, ConfigMap TAB catalog-config, Widget TAB w1, in ascending kind name order.
  6. In /root/ops-backup/victim.yaml, write the ConfigMap orphan-victim — namespace ops-shop-restored, data note: restored-with-stale-owner, and in metadata.ownerReferences, put an owner that does not exist in this cluster (apiVersion: v1, kind: ConfigMap, name: catalog-config-old, and any UUID for uid). After applying it, wait with a condition loop until that object disappears. Then, in /root/ops-backup/fixed.yaml, write a ConfigMap orphan-fixed with the same data without ownerReferences and apply it, and leave the result as two lines in /root/ops-backup/owner.tsv — orphan-victim TAB gone, orphan-fixed TAB alive.
  7. In /root/ops-backup/gadget.yaml, write the custom resource Gadget g1 — apiVersion: ops.example.com/v1, namespace ops-shop-restored, and spec.color is blue. Apply it first while the CRD for that kind does not yet exist, and save the error to /root/ops-backup/order-error.txt (including standard error). Then, in /root/ops-backup/gadget-crd.yaml, write the CRD gadgets.ops.example.com (group ops.example.com, kind Gadget, plural gadgets, namespaced scope, version v1, string spec.color), apply it, and then apply the custom resource again. Finally, in /root/ops-backup/restore-order.txt, write the restore order as four lines — namespace, crd, custom-resource, and workload, one per line, in the correct order.
  8. In /root/ops-backup/inventory.tsv, write the restore list as four lines — <종류> TAB <이름> TAB <기대값> (kind, name, expected value), and the expected value is spec.replicas for the Deployment, spec.ports[0].port for the Service, the number of keys of data for the ConfigMap, and spec.size for the Widget. Then create /root/ops-backup/verify-restore.sh — it reads this table and compares it against the actual objects in ops-shop-restored, and prints OK <종류>/<이름> <값> (kind, name, value) if they match and MISMATCH <종류>/<이름> 기대=<값> 실제=<값> (kind, name, expected value, actual value) if they do not, to standard output only, and it must exit with a non-zero code if even one line is wrong. Save that output to /root/ops-backup/verify.txt.

Reference

Create the set to bring back

In /root/ops-backup/crd.yaml, write the CustomResourceDefinition widgets.ops.example.com — group ops.example.com, kind Widget (plural widgets), namespaced scope, version v1 (served and storage both true), and the schema has an integer spec.size. Then, in /root/ops-backup/scene.yaml, write four objects of the namespace ops-shop in one file — the Deployment catalog (2 replicas, label app: catalog, image nginx:1.27.3), a ClusterIP Service catalog with the same name (selector app: catalog, port 80), the ConfigMap catalog-config (currency: KRW, page-size: "20"), and the Widget w1 (spec.size 3). Create the namespace and apply everything.

Half of a backup lab is deciding what to back up. There is a reason to deliberately create a set with mixed kinds — workloads, configuration, and custom resources each differ in the fields that must be removed on export and in the restore order. A CRD is cluster-scoped, so it remains even if you delete the namespace.

Export it exactly as it is

Export each of the four objects as is, without processing under /root/ops-backup/raw/ — deploy.yaml (Deployment catalog), svc.yaml (Service catalog), cm.yaml (ConfigMap catalog-config), widget.yaml (Widget w1). Save the output of kubectl get <종류> <이름> -o yaml (kind and name) without touching it.

It is important to first look at the raw form. It has everything you need to remove in the next step — identifiers that have meaning only in this cluster, state filled in by controllers, and the history of managed fields. What is a problem and why, you can know only by looking once before removing.

Remove what has meaning only in that cluster

Create /root/ops-backup/clean.sh — it reads one YAML file given as an argument and prints the cleaned result to standard output only. What must be removed is, from metadata, uid, resourceVersion, creationTimestamp, generation, managedFields, ownerReferences, and namespace; the top-level status; and the Service's spec.clusterIP, spec.clusterIPs, and spec.ports[].nodePort. With this script, clean the four files in raw/ and save them under the same names in /root/ops-backup/clean/.

The reason you remove even metadata.namespace is to decide where to restore on the command side — that lets you use the same backup as it is in another namespace or another cluster. yq's del does not raise an error even if you tell it to delete a path that does not exist, so you do not need to split the script by kind.

Create the incident

Delete the namespace ops-shop entirely, and after confirming it is gone, record it on one line in /root/ops-backup/gone.txt — ops-shop TAB NotFound. Before deleting, first check that all four files in /root/ops-backup/raw/ exist.

Deleting a namespace takes everything inside it. A CRD is cluster-scoped and remains, but custom resources created with that CRD disappear together with the namespace. Wait until the deletion is finished and then record it — kubectl delete ns waits for you, but it is safer to check for yourself.

Bring it back into a different namespace

Create the namespace ops-shop-restored and restore the four files in /root/ops-backup/clean/ into that namespace (the cleaned manifests have no namespace, so you specify it in the command). Wait until the Deployment's 2 Pods are Running, then save it as four lines to /root/ops-backup/restored.tsv — Deployment TAB catalog, Service TAB catalog, ConfigMap TAB catalog-config, Widget TAB w1, in ascending kind name order.

Deciding the target namespace in the command is the advantage of object-level backup — you can build a procedure of standing it up once in a verification namespace with the same backup and then doing the real recovery. A new namespace takes a moment for its default service account to appear.

If you leave a stale owner, the restored object disappears

In /root/ops-backup/victim.yaml, write the ConfigMap orphan-victim — namespace ops-shop-restored, data note: restored-with-stale-owner, and in metadata.ownerReferences, put an owner that does not exist in this cluster (apiVersion: v1, kind: ConfigMap, name: catalog-config-old, and any UUID for uid). After applying it, wait with a condition loop until that object disappears. Then, in /root/ops-backup/fixed.yaml, write a ConfigMap orphan-fixed with the same data without ownerReferences and apply it, and leave the result as two lines in /root/ops-backup/owner.tsv — orphan-victim TAB gone, orphan-fixed TAB alive.

If you keep ownerReferences in the backup as they are, the restored object quietly disappears. An owner's uid differs from cluster to cluster, and if the restored cluster cannot find an object with that uid, the garbage collector sees the owner as gone and deletes it. It is an incident whose cause is hard to find because no error or event is left.

If you get the order wrong, the restore fails

In /root/ops-backup/gadget.yaml, write the custom resource Gadget g1 — apiVersion: ops.example.com/v1, namespace ops-shop-restored, and spec.color is blue. Apply it first while the CRD for that kind does not yet exist, and save the error to /root/ops-backup/order-error.txt (including standard error). Then, in /root/ops-backup/gadget-crd.yaml, write the CRD gadgets.ops.example.com (group ops.example.com, kind Gadget, plural gadgets, namespaced scope, version v1, string spec.color), apply it, and then apply the custom resource again. Finally, in /root/ops-backup/restore-order.txt, write the restore order as four lines — namespace, crd, custom-resource, and workload, one per line, in the correct order.

An API server that does not know a kind will not accept the manifest at all. Read as it is what the error sentence says to install first. If you do not write this order down when you make the restore procedure into a document, you will meet the same error on the day of recovery, and then you will not have the leisure you have now.

Stop a person from counting by eye whether the restore worked

In /root/ops-backup/inventory.tsv, write the restore list as four lines — <종류> TAB <이름> TAB <기대값> (kind, name, expected value), and the expected value is spec.replicas for the Deployment, spec.ports[0].port for the Service, the number of keys of data for the ConfigMap, and spec.size for the Widget. Then create /root/ops-backup/verify-restore.sh — it reads this table and compares it against the actual objects in ops-shop-restored, and prints OK <종류>/<이름> <값> (kind, name, value) if they match and MISMATCH <종류>/<이름> 기대=<값> 실제=<값> (kind, name, expected value, actual value) if they do not, to standard output only, and it must exit with a non-zero code if even one line is wrong. Save that output to /root/ops-backup/verify.txt.

If a person looks at the screen after a restore and counts whether everything is there, they will certainly miss one. Moreover, an object existing and its values being the same are different problems — if you only count names, a restore in which the replicas dropped from 2 to 1 also looks like a success. If the script writes a file itself, it overwrites the student's artifact when the grader runs it again, so send it out to standard output only.