Building a Reconcile Loop in Shell
Goal
Implement by yourself, as a shell script, a reconcile loop that reads a CR's spec, brings child resources into line, and writes the result back to status, and prove with the object version that the second run changes nothing.
Why it matters
What you build in this lab is not a controller binary but the essence of the reconcile function. In a real controller too, what is passed to reconcile is only a key, not the object, and it does not tell you what changed. So reconcile always starts over from "what is the desired state now, and what is the actual state." What this design enforces is idempotency. If you write again in an already-correct state, a new watch event is generated, and that event calls reconcile again, triggering itself endlessly. That is why the real test of idempotency is not "no error occurred" but "the child object's resourceVersion stayed the same." Another important distinction is between errors and requeues. A dependency that is not ready yet is not a failure, so you must not return it as an error. Errors pile up exponential backoff and pollute the logs, and that object's retry interval grows longer, so the response is slow when a real problem occurs. Finally, finalizers break the intuition that "delete means vanish immediately." A delete request only stamps a deletion timestamp, and the moment reconcile is called one more time is the last chance to clean up external resources.
Steps
Before you start: lab Pods start fresh for every lab, so the cluster state from the previous lab is not there. If kubectl get crd webservices.apps.labhub.io returns nothing, rewrite and apply the CRD and also run kubectl create ns crd-lab. This lab must have subresources.status and the definitions of replicas, observedGeneration, and conditions under properties.status, and spec needs image (required), replicas (default 1), and tier (default dev).
- First, create a WebService
checkoutincrd-lab(spec.image: nginx:1.27,spec.replicas: 3,spec.tier: dev). Then create/root/op/reconcile/reconcile.shand runchmod +xon it. This script must usekubectl getto read the CR's spec (desired) and the child ConfigMapcheckout-desired(actual) and save them to/root/op/reconcile/out/observe.jsonin the form{"desired": {...}, "actual": {...}}, anddesired.imagemust be exactly equal to the CR'sspec.image. - Save the judgment of the first run to
/root/op/reconcile/out/decision-1.json.actioniscreate,reasonis a string describing why it judged that way, andtargetholds what the judgment is about (for example,configmap/checkout-desired). - Make the script carry out the judgment. Create a ConfigMap
checkout-desiredincrd-labwheredata.imageis the CR'sspec.imagevalue,metadata.ownerReferences[0].uidis the actual uid ofcheckout, andmetadata.labelshasapp.kubernetes.io/managed-by: webservice-controller. - Save the value of
kubectl get cm checkout-desired -n crd-lab -o jsonpath='{.metadata.resourceVersion}'to/root/op/reconcile/out/rv-before.txt, run the script once more, and then save the same value to/root/op/reconcile/out/rv-after.txt. The two values must be the same, and the second judgment must be recorded in/root/op/reconcile/out/decision-2.jsonwithactionset tonoop. - Make the script write status back. Put a condition with
type: Readyandstatus: "True"incheckout'sstatus.conditions, and makestatus.observedGenerationequal tometadata.generationandstatus.replicasequal tospec.replicas. The script must actually contain the string--subresource=status. - Create
/root/op/reconcile/out/requeue.json.actionisrequeue,after_secondsis an integer greater than 0, andis_errorisfalse. Then in/root/op/reconcile/out/backoff-note.txt, write two things, each in at least one sentence — why backoff, where the retry interval grows exponentially, is needed, and the problem that if you treat every "not ready yet" as an error, error logs explode and one object starves the queue. - Create a WebService
ephemeralincrd-laband putwebservice.labhub.io/cleanupinmetadata.finalizers. Also create a child ConfigMapephemeral-desiredin the same way as in step 3. Then request deletion withkubectl delete webservice ephemeral -n crd-lab --wait=false, and save the whole object right after that to/root/op/reconcile/out/terminating.json(this file must show bothmetadata.deletionTimestampandmetadata.finalizers). Then delete the child ConfigMapephemeral-desired, write what you cleaned up to/root/op/reconcile/out/cleanup.txt, and remove the finalizer so thatephemeralactually disappears. - Create two more WebServices,
billingandsearch, incrd-lab(each with an image and replicas), run all three,checkout,billing, andsearch, through the reconciler for one full pass, and then create/root/op/reconcile/out/reconcile-report.json.itemsis an array whose elements each havenameandaction(one ofcreate/update/noop/requeue),summary.noopholds the count of cases where nothing was done, andconvergedholdstrue. For all three CRs,status.observedGenerationmust equalmetadata.generation.
Notes
- Lab Pods start fresh for every lab, so the cluster state from the previous lab is not there. Still, if you leave your declarations as files, you can rebuild the same state on any Pod — this is the practical advantage of the declarative approach.
- Removing a finalizer:
kubectl patch webservice ephemeral -n crd-lab --type=merge -p '{"metadata":{"finalizers":null}}' - If you run
kubectl applywith identical content, the server treats it as no change andresourceVersiondoes not rise. However, if you put a value that changes every time (such as a timestamp) into an annotation, that property breaks. - Write status with a patch:
kubectl patch webservice checkout -n crd-lab --subresource=status --type=merge -p '{"status":{...}}' - Common mistake 1: having the script do
kubectl replace, or a forced apply afterkubectl create --dry-run, every time in step 4. Even with identical content the version rises, and the idempotency check fails. - Common mistake 2: writing
is_errorastruein step 6. A dependency that does not exist yet is not an error but a waiting situation. - Common mistake 3: waiting for the delete command to finish in step 7. While a finalizer is attached the command does not return, so you must use the option that does not wait in order to observe the terminating state.
Read the desired state and the actual state
First, create a WebService checkout in crd-lab (spec.image: nginx:1.27, spec.replicas: 3, spec.tier: dev). Then create /root/op/reconcile/reconcile.sh and run chmod +x on it. This script must use kubectl get to read the CR's spec (desired) and the child ConfigMap checkout-desired (actual) and save them to /root/op/reconcile/out/observe.json in the form {"desired": {...}, "actual": {...}}, and desired.image must be exactly equal to the CR's spec.image.
The first step of reconcile is not judging but reading. The CR's spec is desired, and the child object is actual. If the child object does not exist yet, it is normal for actual to be empty, and the script needs execute permission.
Compare and make the reconcile judgment
Save the judgment of the first run to /root/op/reconcile/out/decision-1.json. action is create, reason is a string describing why it judged that way, and target holds what the judgment is about (for example, configmap/checkout-desired).
A judgment must not end with just an action. You must also leave why it judged that way and what it is about, so that you can trace it later from the logs alone.
Create the child resource as judged
Make the script carry out the judgment. Create a ConfigMap checkout-desired in crd-lab where data.image is the CR's spec.image value, metadata.ownerReferences[0].uid is the actual uid of checkout, and metadata.labels has app.kubernetes.io/managed-by: webservice-controller.
Do not copy the values by hand; read them from the CR and put them in. You need both the link to the parent and the marker label that shows you created it.
Make the second run change nothing
Save the value of kubectl get cm checkout-desired -n crd-lab -o jsonpath='{.metadata.resourceVersion}' to /root/op/reconcile/out/rv-before.txt, run the script once more, and then save the same value to /root/op/reconcile/out/rv-after.txt. The two values must be the same, and the second judgment must be recorded in /root/op/reconcile/out/decision-2.json with action set to noop.
The evidence of idempotency is not the log but the object's version. Record the child object's resourceVersion before and after the second run and compare them. If the content is the same, the server does not write again.
Write the reconcile result back to status
Make the script write status back. Put a condition with type: Ready and status: "True" in checkout's status.conditions, and make status.observedGeneration equal to metadata.generation and status.replicas equal to spec.replicas. The script must actually contain the string --subresource=status.
You write status through a different path from spec. The string that specifies that path must actually be in the script. Match the processed generation number and the observed replica count too.
Sort out the requeue decision and backoff
Create /root/op/reconcile/out/requeue.json. action is requeue, after_seconds is an integer greater than 0, and is_error is false. Then in /root/op/reconcile/out/backoff-note.txt, write two things, each in at least one sentence — why backoff, where the retry interval grows exponentially, is needed, and the problem that if you treat every "not ready yet" as an error, error logs explode and one object starves the queue.
A dependency that does not exist yet is not a failure. Write in seconds when to look again, and state explicitly that this is not an error. In the note, write both why the retry interval grows and what problem arises if you treat everything as an error.
Clean up and delete through a finalizer
Create a WebService ephemeral in crd-lab and put webservice.labhub.io/cleanup in metadata.finalizers. Also create a child ConfigMap ephemeral-desired in the same way as in step 3. Then request deletion with kubectl delete webservice ephemeral -n crd-lab --wait=false, and save the whole object right after that to /root/op/reconcile/out/terminating.json (this file must show both metadata.deletionTimestamp and metadata.finalizers). Then delete the child ConfigMap ephemeral-desired, write what you cleaned up to /root/op/reconcile/out/cleanup.txt, and remove the finalizer so that ephemeral actually disappears.
A delete request is not an immediate deletion. First save how the object looks right after the deletion timestamp is stamped, do the cleanup work, and then remove the finalizer, and only then does the object disappear. There is a delete option that does not wait.
Run several CRs through one pass and build a report
Create two more WebServices, billing and search, in crd-lab (each with an image and replicas), run all three, checkout, billing, and search, through the reconciler for one full pass, and then create /root/op/reconcile/out/reconcile-report.json. items is an array whose elements each have name and action (one of create/update/noop/requeue), summary.noop holds the count of cases where nothing was done, and converged holds true. For all three CRs, status.observedGeneration must equal metadata.generation.
The reconciler is not dedicated to one particular CR. Loop over several targets, collect each judgment, and summarize in one line whether all of them reached the desired state. It has converged only when every target's generation number matches.