TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

You deleted the parent and the children stayed - owner references and deletion propagation

Continue in TT Lab

Goal

Create for yourself what each of the four fields of an owner reference decides, prove the difference among the three --cascade modes with objects, and then build a report that diagnoses an object stuck in Terminating.

Why it matters

Resources an Operator creates do not live alone. When one parent disappears, the things attached below it must be cleaned up together, but Kubernetes has no reference count to do this. Instead there is a single ownerReferences line in which a child points to its parent, and the garbage collector cleans up by following that reference. So writing the reference wrongly causes a quiet failure — if the uid is wrong, a child you just created disappears within seconds, and if it crosses a namespace, the child is cleaned up leaving only one warning event. Conversely, a finalizer is a device that holds deletion until cleanup is finished, so if there is no controller to do that cleanup, the object stays in Terminating forever. The screens operators actually run into are usually one of these two, and both need a procedure for finding the cause.

Steps

  1. In /root/op-ownership/site-crd.yaml, write the CRD sites.own.labhub.io — group own.labhub.io, kind Site, plural sites, a single version v1, with only spec.region (string) in the schema. Create the namespace op-own, apply Site alpha (region kr-east) with /root/op-ownership/site-alpha.yaml, and then save that object's metadata.uid to /root/op-ownership/alpha-uid.txt on one line.
  2. In /root/op-ownership/alpha-good.yaml, write a ConfigMap alpha-good (namespace op-own) and put Site alpha in ownerReferences — apiVersion own.labhub.io/v1, kind Site, name alpha, and for uid the actual value you got in step 1. In /root/op-ownership/alpha-bad.yaml, write a ConfigMap alpha-bad with the same content but with only the uid written as 00000000-0000-0000-0000-000000000000. Apply both, wait a moment, and then save the output of kubectl -n op-own get cm to /root/op-ownership/gc-uid.txt.
  3. Create the namespace op-own-remote and write a ConfigMap alpha-remote in /root/op-ownership/alpha-remote.yaml — its namespace is op-own-remote, but its ownerReferences points (with the actual uid) to Site alpha in op-own. After applying, wait a moment and then save the output of kubectl -n op-own-remote get events to /root/op-ownership/xns-event.txt.
  4. Create Site beta (region kr-west) with /root/op-ownership/site-beta.yaml, and as its children put ConfigMaps beta-1 and beta-2 in one file, /root/op-ownership/beta-children.yaml (join the two with ---), set owner references with the actual uid, and apply it. Then delete with kubectl -n op-own delete site beta --cascade=background, wait until the children disappear, and save the output of kubectl -n op-own get cm to /root/op-ownership/cascade-background.txt.
  5. Create Site gamma (region jp-east) with /root/op-ownership/site-gamma.yaml, put the child ConfigMaps gamma-1 and gamma-2 in /root/op-ownership/gamma-children.yaml, link them with the actual uid, and apply. After deleting with kubectl -n op-own delete site gamma --cascade=orphan, save the result of kubectl -n op-own get cm gamma-1 -o jsonpath='{.metadata.ownerReferences}' to /root/op-ownership/cascade-orphan.txt as one line, gamma-1-ownerrefs=<값, 비어 있으면 none> (the value, or none if empty).
  6. Create Site delta (region us-west) with /root/op-ownership/site-delta.yaml, and create a ConfigMap delta-1 with /root/op-ownership/delta-child.yaml — put blockOwnerDeletion: true together with the actual uid in the owner reference, and attach the finalizer own.labhub.io/hold to this ConfigMap itself. Then run kubectl -n op-own delete site delta --cascade=foreground --wait=false, and after a moment save kubectl -n op-own get site delta -o json to /root/op-ownership/foreground-stuck.json. Leave this Site as it is until this lab ends.
  7. Create Site epsilon (region eu-west) with /root/op-ownership/site-epsilon.yaml together with the finalizer own.labhub.io/drain, link the child ConfigMap epsilon-data with the actual uid, and apply it with /root/op-ownership/epsilon-child.yaml. Save the object right after running kubectl -n op-own delete site epsilon --wait=false to /root/op-ownership/finalizer-pending.json, do the cleanup work (deleting the child ConfigMap directly), and then write at least one line in /root/op-ownership/epsilon-cleanup.txt about what you cleaned up. Finally, remove the finalizer so that the Site actually disappears.
  8. Create /root/op-ownership/stuck-report.sh — in op-own, print (1) any Site with deletionTimestamp stamped as STUCK Site/<이름> finalizers=<쉼표로 이은 목록> (name, and a comma-separated list of finalizers), and (2) any ConfigMap with an owner reference whose blockOwnerDeletion is true as BLOCKER ConfigMap/<이름> owner=<종류>/<이름> finalizers=<목록> (name, the owner's kind and name, and the finalizer list), merge the two kinds, sort them, and emit only to standard output. If there is even one STUCK line, it must exit with a nonzero code, and if not, print NONE and exit with 0. Save the output to /root/op-ownership/stuck-report.txt, and write in /root/op-ownership/diagnosis.txt, in sentences a person can read, who is blocking what and why (the name of the blocking child and the name of its finalizer must be included).

Notes

Create the type that will be the parent

In /root/op-ownership/site-crd.yaml, write the CRD sites.own.labhub.io — group own.labhub.io, kind Site, plural sites, a single version v1, with only spec.region (string) in the schema. Create the namespace op-own, apply Site alpha (region kr-east) with /root/op-ownership/site-alpha.yaml, and then save that object's metadata.uid to /root/op-ownership/alpha-uid.txt on one line.

In an owner reference, the name is a value for humans to read, and the real key is the uid. If an object with the same name is deleted and recreated, the uid changes, and this property becomes the crux of the steps that follow. Pull out the uid with kubectl get … -o jsonpath.

If you get the uid wrong, the garbage collector deletes the child

In /root/op-ownership/alpha-good.yaml, write a ConfigMap alpha-good (namespace op-own) and put Site alpha in ownerReferences — apiVersion own.labhub.io/v1, kind Site, name alpha, and for uid the actual value you got in step 1. In /root/op-ownership/alpha-bad.yaml, write a ConfigMap alpha-bad with the same content but with only the uid written as 00000000-0000-0000-0000-000000000000. Apply both, wait a moment, and then save the output of kubectl -n op-own get cm to /root/op-ownership/gc-uid.txt.

The garbage collector finds the parent by uid, not by the reference's name. If it cannot find one, it sees it as "a child whose parent has already disappeared" and cleans it up. Because of this behavior, if a controller caches a uid and uses it after the parent has been recreated, a child you just created disappears right away. When waiting, a short loop that runs until the object disappears is safer than a fixed sleep.

There is no ownership across namespaces

Create the namespace op-own-remote and write a ConfigMap alpha-remote in /root/op-ownership/alpha-remote.yaml — its namespace is op-own-remote, but its ownerReferences points (with the actual uid) to Site alpha in op-own. After applying, wait a moment and then save the output of kubectl -n op-own-remote get events to /root/op-ownership/xns-event.txt.

An owner reference is formed only within the same namespace, and only a cluster-scoped object can have namespaced children. If you break the rule, the API server does not block the apply — the garbage collector handles it later and leaves a trace as an event. Read the event's reason (REASON) word exactly as it is.

background deletes the parent first

Create Site beta (region kr-west) with /root/op-ownership/site-beta.yaml, and as its children put ConfigMaps beta-1 and beta-2 in one file, /root/op-ownership/beta-children.yaml (join the two with ---), set owner references with the actual uid, and apply it. Then delete with kubectl -n op-own delete site beta --cascade=background, wait until the children disappear, and save the output of kubectl -n op-own get cm to /root/op-ownership/cascade-background.txt.

background is the default. The API server deletes the parent immediately and leaves the child cleanup to the garbage collector. So the command returns quickly while the children remain for a moment — this time gap is the source of the illusion "I deleted it but it still shows."

orphan leaves the children and cuts only the tie

Create Site gamma (region jp-east) with /root/op-ownership/site-gamma.yaml, put the child ConfigMaps gamma-1 and gamma-2 in /root/op-ownership/gamma-children.yaml, link them with the actual uid, and apply. After deleting with kubectl -n op-own delete site gamma --cascade=orphan, save the result of kubectl -n op-own get cm gamma-1 -o jsonpath='{.metadata.ownerReferences}' to /root/op-ownership/cascade-orphan.txt as one line, gamma-1-ownerrefs=<값, 비어 있으면 none> (the value, or none if empty).

orphan does not delete the children. Instead, the garbage collector detaches that owner reference from the children — if it did not detach it, a reference with no parent would remain and the child would disappear at the next cleanup. It is the approach you use when you want to remove an Operator while keeping its resources alive.

foreground waits for the children and stalls

Create Site delta (region us-west) with /root/op-ownership/site-delta.yaml, and create a ConfigMap delta-1 with /root/op-ownership/delta-child.yaml — put blockOwnerDeletion: true together with the actual uid in the owner reference, and attach the finalizer own.labhub.io/hold to this ConfigMap itself. Then run kubectl -n op-own delete site delta --cascade=foreground --wait=false, and after a moment save kubectl -n op-own get site delta -o json to /root/op-ownership/foreground-stuck.json. Leave this Site as it is until this lab ends.

foreground reverses the order — it deletes the parent only after all the children are gone. To express that waiting, the API server attaches a finalizer to the parent, and you can find its name in the JSON you saved. If a child has a finalizer, that child cannot disappear, so the wait never ends.

Deletion does not finish until you remove the finalizer

Create Site epsilon (region eu-west) with /root/op-ownership/site-epsilon.yaml together with the finalizer own.labhub.io/drain, link the child ConfigMap epsilon-data with the actual uid, and apply it with /root/op-ownership/epsilon-child.yaml. Save the object right after running kubectl -n op-own delete site epsilon --wait=false to /root/op-ownership/finalizer-pending.json, do the cleanup work (deleting the child ConfigMap directly), and then write at least one line in /root/op-ownership/epsilon-cleanup.txt about what you cleaned up. Finally, remove the finalizer so that the Site actually disappears.

If a finalizer is attached, the delete request only stamps deletionTimestamp and stops. The object can still be queried and modified, but you cannot create a new one. What a controller does is exactly this interval — it cleans up the outside system and then removes its own finalizer. To remove one element from a list, the remove operation of kubectl patch --type=json is convenient.

Diagnose what is stuck in Terminating

Create /root/op-ownership/stuck-report.sh — in op-own, print (1) any Site with deletionTimestamp stamped as STUCK Site/<이름> finalizers=<쉼표로 이은 목록> (name, and a comma-separated list of finalizers), and (2) any ConfigMap with an owner reference whose blockOwnerDeletion is true as BLOCKER ConfigMap/<이름> owner=<종류>/<이름> finalizers=<목록> (name, the owner's kind and name, and the finalizer list), merge the two kinds, sort them, and emit only to standard output. If there is even one STUCK line, it must exit with a nonzero code, and if not, print NONE and exit with 0. Save the output to /root/op-ownership/stuck-report.txt, and write in /root/op-ownership/diagnosis.txt, in sentences a person can read, who is blocking what and why (the name of the blocking child and the name of its finalizer must be included).

The key to the diagnosis is putting two lists side by side — the stalled parent, and the children that can hold that parent. Do not put values that change every time you look, such as age or timestamps, in the output. That way you can keep this report as a file and compare it with later results.