CNPA — Cloud Native Platform Engineering Associate
The quota we deleted came back seconds later
Goal
On a real k3s, you attach a small controller to a platform API called TeamSpace. When a user creates a single TeamSpace, a namespace and a quota follow automatically, they revert even when someone edits them by hand, and on deletion the object disappears only after the cleanup work is done — and you confirm each of these with values at every step.
Why it matters
A CRD only registers a new noun with the API server, and it is the controller that gives that noun meaning. A Kubernetes controller is a reconciliation loop that endlessly compares the desired state (spec) with the actual state and narrows the difference, and the Operator pattern is a way of putting a particular domain's operational knowledge into that loop. When a platform team builds a self-service API, following this structure means the API server provides storage, authorization, auditing, and watch, and the team only has to write the reconciliation logic. Because reconciliation runs on the current state rather than on events, it converges eventually even after a missed notification or a controller restart, and users read "has my change been applied?" from the status field observedGeneration. In this lab you break that property one piece at a time to confirm it.
Steps
- Write the CRD
teamspaces.platform.labhub.ioin/root/cnpa-op/crd.yamland apply it. Use groupplatform.labhub.io,scope: Cluster, kindTeamSpace(pluralteamspaces), enablesubresources.statuson versionv1alpha1, and use this schema:spec.pods(integer, required, 1..50),spec.cpu(string, default"1"),status.observedGeneration(integer), andstatus.namespace(string). Then create the TeamSpacealpha(spec.pods: 4) from/root/cnpa-op/alpha.yaml, and writeuid(the uid of alpha) andnamespace_exists(whether the namespaceteam-alphaexists, a boolean) into/root/cnpa-op/before.json. - Do three things to alpha in turn, and read
metadata.generationeach time. ① Add the labelowner=platform, ② apply a merge patch of{"status":{"observedGeneration":7}}to the main resource address, ③ changespec.podsto6. Writeuid,gen_initial(the starting value),gen_after_label,gen_after_spec(after ③), andstatus_via_main(the value of alpha's.statusright after ②, or null if there is none) into/root/cnpa-op/generation.json. - Write
/root/cnpa-op/reconcile.sh(executable). When run once, for every TeamSpace<이름>(its name) it must bring the namespaceteam-<이름>(labelplatform.labhub.io/teamspace: <이름>, and an ownerReference pointing to that TeamSpace) and the ResourceQuotateam-quotainside it (podsfrom spec.pods,requests.cpufrom spec.cpu) into the desired state, and recordobservedGeneration(the current generation) andnamespacein the status subresource. After running it once, writequota_uid(the uid of team-alpha's team-quota),namespace_owner_uid(the ownerReference uid of team-alpha), andsecond_pass_rv_changed(whether the quota's resourceVersion changed when you ran it one more time, a boolean) into/root/cnpa-op/reconcile.json. - Create
/etc/systemd/system/cnpa-op.serviceso that it runsreconcile.shendlessly at 3-second intervals (Restart=always), then enable and start it. When you create a new TeamSpace, the namespace and quota must appear within a few seconds without anyone doing anything. - While the controller is running, delete the
team-quotainteam-alphaby hand, and measure the seconds it takes to come back with a new uid. Next, patchpodsof the restored quota to99and watch whether it returns to6. Writedeleted_uid,restored_uid,restore_seconds(an integer),edited_to(99), andreverted_to(the value it returned to, an integer) into/root/cnpa-op/drift.json. - Modify
reconcile.shso that it handles the finalizerplatform.labhub.io/cleanup. Attach this finalizer to a TeamSpace that has no deletion request (preserving other finalizers), and for a TeamSpace that has adeletionTimestampset, leavename,uid, andpodsin/root/cnpa-op/archive/<이름>.json(where the file name is the TeamSpace name) and then remove the finalizer. Then create the TeamSpacebeta(pods: 2,cpu: "500m"), see that its namespace appears, and delete it. Writeuid(the uid of beta),blocked_seen(whether you saw beta still present with a deletionTimestamp right after deletion, a boolean), andnamespace_gone(whether team-beta disappeared afterward, a boolean) into/root/cnpa-op/cleanup.json. - Stop
cnpa-op.serviceand change alpha'sspec.podsto8. After waiting 8 seconds, read alpha'smetadata.generationandstatus.observedGenerationand the quota'spods, then restart the service and wait until the two match. Writegeneration,observed_while_stopped,quota_pods_while_stopped,observed_after_start, andquota_pods_after_start(keep the quota values as strings) into/root/cnpa-op/catchup.json. - Write
crd_alone_created_namespace(step 1),status_bumps_generation(whether writing status raised the generation — check in step 2 or on alpha now),restore_seconds(step 5),finalizer(the name),archived(the uid written in beta.json in the archive directory),generations_behind_while_stopped(the difference between the generation in step 7 and observed_while_stopped), andtrigger(which ofleveloredgethis controller follows) into/root/cnpa-op/report.json.
Notes
- k3s is inside the VM, and you can use
kubectl,jq,python3, andsystemctl. Container images are not used. - Writing status:
kubectl patch teamspace <이름> --subresource=status --type merge -p '{"status":{...}}'. - Common mistake: patching status through the main resource address. When a status subresource exists, the server silently discards it and only prints
patched (no change). - Common mistake: a patch that attaches a finalizer overwrites the whole existing list. Other controllers' finalizers disappear.
- Common mistake: ending the lab with the controller stopped. The graders for steps 4, 5, 6, and 7 pass only if the controller is running.
- Controllers · Operator pattern · CRD status subresource · Finalizers · Owners and Dependents · CNCF Platforms White Paper
Registering a type alone makes nothing happen
Write the CRD teamspaces.platform.labhub.io in /root/cnpa-op/crd.yaml and apply it. Use group platform.labhub.io, scope: Cluster, kind TeamSpace (plural teamspaces), enable subresources.status on version v1alpha1, and use this schema: spec.pods (integer, required, 1..50), spec.cpu (string, default "1"), status.observedGeneration (integer), and status.namespace (string). Then create the TeamSpace alpha (spec.pods: 4) from /root/cnpa-op/alpha.yaml, and write uid (the uid of alpha) and namespace_exists (whether the namespace team-alpha exists, a boolean) into /root/cnpa-op/before.json.
A CRD only tells the API server about a new noun. Nothing is yet watching that noun and creating anything. You can tell whether a namespace exists from the exit code of kubectl get ns.
What raises the generation
Do three things to alpha in turn, and read metadata.generation each time. ① Add the label owner=platform, ② apply a merge patch of {"status":{"observedGeneration":7}} to the main resource address, ③ change spec.pods to 6. Write uid, gen_initial (the starting value), gen_after_label, gen_after_spec (after ③), and status_via_main (the value of alpha's .status right after ②, or null if there is none) into /root/cnpa-op/generation.json.
The generation is a number that the API server raises only when the spec changes. If only metadata changes, the resourceVersion goes up but the generation stays the same. On a CRD with the status subresource enabled, the server discards a status change sent to the main resource. Look at what kubectl prints in that case.
The same result no matter how many times it runs
Write /root/cnpa-op/reconcile.sh (executable). When run once, for every TeamSpace <이름> (its name) it must bring the namespace team-<이름> (label platform.labhub.io/teamspace: <이름>, and an ownerReference pointing to that TeamSpace) and the ResourceQuota team-quota inside it (pods from spec.pods, requests.cpu from spec.cpu) into the desired state, and record observedGeneration (the current generation) and namespace in the status subresource. After running it once, write quota_uid (the uid of team-alpha's team-quota), namespace_owner_uid (the ownerReference uid of team-alpha), and second_pass_rv_changed (whether the quota's resourceVersion changed when you ran it one more time, a boolean) into /root/cnpa-op/reconcile.json.
Reconciliation does not look at "what changed" but at "is the desired state the same as the actual state right now." So the result must be the same no matter how many times it runs. kubectl apply does not write if the content is the same. Write status with kubectl patch --subresource=status. The grader runs this script two more times and checks that nothing is written.
It keeps running, not just once
Create /etc/systemd/system/cnpa-op.service so that it runs reconcile.sh endlessly at 3-second intervals (Restart=always), then enable and start it. When you create a new TeamSpace, the namespace and quota must appear within a few seconds without anyone doing anything.
Put a while true; do ...; sleep 3; done loop into ExecStart with bash -c. The grader creates a temporary TeamSpace, checks whether the namespace appears on its own, and deletes it.
The deleted quota came back a few seconds later
While the controller is running, delete the team-quota in team-alpha by hand, and measure the seconds it takes to come back with a new uid. Next, patch pods of the restored quota to 99 and watch whether it returns to 6. Write deleted_uid, restored_uid, restore_seconds (an integer), edited_to (99), and reverted_to (the value it returned to, an integer) into /root/cnpa-op/drift.json.
The controller does not know who deleted it. On the next pass it simply compares the desired state with the actual state. So the restored quota has the same name but is a new object. Recall what you changed spec.pods to in step 2.
There is work to do before deleting
Modify reconcile.sh so that it handles the finalizer platform.labhub.io/cleanup. Attach this finalizer to a TeamSpace that has no deletion request (preserving other finalizers), and for a TeamSpace that has a deletionTimestamp set, leave name, uid, and pods in /root/cnpa-op/archive/<이름>.json (where the file name is the TeamSpace name) and then remove the finalizer. Then create the TeamSpace beta (pods: 2, cpu: "500m"), see that its namespace appears, and delete it. Write uid (the uid of beta), blocked_seen (whether you saw beta still present with a deletionTimestamp right after deletion, a boolean), and namespace_gone (whether team-beta disappeared afterward, a boolean) into /root/cnpa-op/cleanup.json.
If you send a delete to an object that has a finalizer, the API server does not remove it and only sets the deletionTimestamp. It actually disappears only when the list is empty. The namespace is removed by the garbage collector following the ownerReference, so the controller does not need to delete it directly. Use a finalizer only for things that cannot be expressed as an ownership relationship, such as a record outside the cluster. To see the moment right after deletion, send only the request with kubectl delete --wait=false and read it immediately. The grader creates one more temporary TeamSpace and deletes it to test.
Don't miss changes made while stopped
Stop cnpa-op.service and change alpha's spec.pods to 8. After waiting 8 seconds, read alpha's metadata.generation and status.observedGeneration and the quota's pods, then restart the service and wait until the two match. Write generation, observed_while_stopped, quota_pods_while_stopped, observed_after_start, and quota_pods_after_start (keep the quota values as strings) into /root/cnpa-op/catchup.json.
This controller does not receive and handle change events. On every pass it simply reads the current state, so it does not need to know what changed, or how many times, while it was stopped. An observedGeneration smaller than the generation is a signal that "this spec has not been applied yet," and it is the standard way for users to read a controller's progress.
Record what the controller did as values
Write crd_alone_created_namespace (step 1), status_bumps_generation (whether writing status raised the generation — check in step 2 or on alpha now), restore_seconds (step 5), finalizer (the name), archived (the uid written in beta.json in the archive directory), generations_behind_while_stopped (the difference between the generation in step 7 and observed_while_stopped), and trigger (which of level or edge this controller follows) into /root/cnpa-op/report.json.
Calculate from the JSON you left in the earlier steps. This controller does not react to each individual change event; every time it reads and reconciles the entire current state. The grader checks the same files against the cluster again.