TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

Building a Reconcile Loop in Shell

Continue in TT Lab

Goal

Implement by yourself, as a shell script, a reconcile loop that reads a CR's spec, brings child resources into line, and writes the result back to status, and prove with the object version that the second run changes nothing.

Why it matters

What you build in this lab is not a controller binary but the essence of the reconcile function. In a real controller too, what is passed to reconcile is only a key, not the object, and it does not tell you what changed. So reconcile always starts over from "what is the desired state now, and what is the actual state." What this design enforces is idempotency. If you write again in an already-correct state, a new watch event is generated, and that event calls reconcile again, triggering itself endlessly. That is why the real test of idempotency is not "no error occurred" but "the child object's resourceVersion stayed the same." Another important distinction is between errors and requeues. A dependency that is not ready yet is not a failure, so you must not return it as an error. Errors pile up exponential backoff and pollute the logs, and that object's retry interval grows longer, so the response is slow when a real problem occurs. Finally, finalizers break the intuition that "delete means vanish immediately." A delete request only stamps a deletion timestamp, and the moment reconcile is called one more time is the last chance to clean up external resources.

Steps

Before you start: lab Pods start fresh for every lab, so the cluster state from the previous lab is not there. If kubectl get crd webservices.apps.labhub.io returns nothing, rewrite and apply the CRD and also run kubectl create ns crd-lab. This lab must have subresources.status and the definitions of replicas, observedGeneration, and conditions under properties.status, and spec needs image (required), replicas (default 1), and tier (default dev).

  1. First, create a WebService checkout in crd-lab (spec.image: nginx:1.27, spec.replicas: 3, spec.tier: dev). Then create /root/op/reconcile/reconcile.sh and run chmod +x on it. This script must use kubectl get to read the CR's spec (desired) and the child ConfigMap checkout-desired (actual) and save them to /root/op/reconcile/out/observe.json in the form {"desired": {...}, "actual": {...}}, and desired.image must be exactly equal to the CR's spec.image.
  2. Save the judgment of the first run to /root/op/reconcile/out/decision-1.json. action is create, reason is a string describing why it judged that way, and target holds what the judgment is about (for example, configmap/checkout-desired).
  3. Make the script carry out the judgment. Create a ConfigMap checkout-desired in crd-lab where data.image is the CR's spec.image value, metadata.ownerReferences[0].uid is the actual uid of checkout, and metadata.labels has app.kubernetes.io/managed-by: webservice-controller.
  4. Save the value of kubectl get cm checkout-desired -n crd-lab -o jsonpath='{.metadata.resourceVersion}' to /root/op/reconcile/out/rv-before.txt, run the script once more, and then save the same value to /root/op/reconcile/out/rv-after.txt. The two values must be the same, and the second judgment must be recorded in /root/op/reconcile/out/decision-2.json with action set to noop.
  5. Make the script write status back. Put a condition with type: Ready and status: "True" in checkout's status.conditions, and make status.observedGeneration equal to metadata.generation and status.replicas equal to spec.replicas. The script must actually contain the string --subresource=status.
  6. Create /root/op/reconcile/out/requeue.json. action is requeue, after_seconds is an integer greater than 0, and is_error is false. Then in /root/op/reconcile/out/backoff-note.txt, write two things, each in at least one sentence — why backoff, where the retry interval grows exponentially, is needed, and the problem that if you treat every "not ready yet" as an error, error logs explode and one object starves the queue.
  7. Create a WebService ephemeral in crd-lab and put webservice.labhub.io/cleanup in metadata.finalizers. Also create a child ConfigMap ephemeral-desired in the same way as in step 3. Then request deletion with kubectl delete webservice ephemeral -n crd-lab --wait=false, and save the whole object right after that to /root/op/reconcile/out/terminating.json (this file must show both metadata.deletionTimestamp and metadata.finalizers). Then delete the child ConfigMap ephemeral-desired, write what you cleaned up to /root/op/reconcile/out/cleanup.txt, and remove the finalizer so that ephemeral actually disappears.
  8. Create two more WebServices, billing and search, in crd-lab (each with an image and replicas), run all three, checkout, billing, and search, through the reconciler for one full pass, and then create /root/op/reconcile/out/reconcile-report.json. items is an array whose elements each have name and action (one of create/update/noop/requeue), summary.noop holds the count of cases where nothing was done, and converged holds true. For all three CRs, status.observedGeneration must equal metadata.generation.

Notes

Read the desired state and the actual state

First, create a WebService checkout in crd-lab (spec.image: nginx:1.27, spec.replicas: 3, spec.tier: dev). Then create /root/op/reconcile/reconcile.sh and run chmod +x on it. This script must use kubectl get to read the CR's spec (desired) and the child ConfigMap checkout-desired (actual) and save them to /root/op/reconcile/out/observe.json in the form {"desired": {...}, "actual": {...}}, and desired.image must be exactly equal to the CR's spec.image.

The first step of reconcile is not judging but reading. The CR's spec is desired, and the child object is actual. If the child object does not exist yet, it is normal for actual to be empty, and the script needs execute permission.

Compare and make the reconcile judgment

Save the judgment of the first run to /root/op/reconcile/out/decision-1.json. action is create, reason is a string describing why it judged that way, and target holds what the judgment is about (for example, configmap/checkout-desired).

A judgment must not end with just an action. You must also leave why it judged that way and what it is about, so that you can trace it later from the logs alone.

Create the child resource as judged

Make the script carry out the judgment. Create a ConfigMap checkout-desired in crd-lab where data.image is the CR's spec.image value, metadata.ownerReferences[0].uid is the actual uid of checkout, and metadata.labels has app.kubernetes.io/managed-by: webservice-controller.

Do not copy the values by hand; read them from the CR and put them in. You need both the link to the parent and the marker label that shows you created it.

Make the second run change nothing

Save the value of kubectl get cm checkout-desired -n crd-lab -o jsonpath='{.metadata.resourceVersion}' to /root/op/reconcile/out/rv-before.txt, run the script once more, and then save the same value to /root/op/reconcile/out/rv-after.txt. The two values must be the same, and the second judgment must be recorded in /root/op/reconcile/out/decision-2.json with action set to noop.

The evidence of idempotency is not the log but the object's version. Record the child object's resourceVersion before and after the second run and compare them. If the content is the same, the server does not write again.

Write the reconcile result back to status

Make the script write status back. Put a condition with type: Ready and status: "True" in checkout's status.conditions, and make status.observedGeneration equal to metadata.generation and status.replicas equal to spec.replicas. The script must actually contain the string --subresource=status.

You write status through a different path from spec. The string that specifies that path must actually be in the script. Match the processed generation number and the observed replica count too.

Sort out the requeue decision and backoff

Create /root/op/reconcile/out/requeue.json. action is requeue, after_seconds is an integer greater than 0, and is_error is false. Then in /root/op/reconcile/out/backoff-note.txt, write two things, each in at least one sentence — why backoff, where the retry interval grows exponentially, is needed, and the problem that if you treat every "not ready yet" as an error, error logs explode and one object starves the queue.

A dependency that does not exist yet is not a failure. Write in seconds when to look again, and state explicitly that this is not an error. In the note, write both why the retry interval grows and what problem arises if you treat everything as an error.

Clean up and delete through a finalizer

Create a WebService ephemeral in crd-lab and put webservice.labhub.io/cleanup in metadata.finalizers. Also create a child ConfigMap ephemeral-desired in the same way as in step 3. Then request deletion with kubectl delete webservice ephemeral -n crd-lab --wait=false, and save the whole object right after that to /root/op/reconcile/out/terminating.json (this file must show both metadata.deletionTimestamp and metadata.finalizers). Then delete the child ConfigMap ephemeral-desired, write what you cleaned up to /root/op/reconcile/out/cleanup.txt, and remove the finalizer so that ephemeral actually disappears.

A delete request is not an immediate deletion. First save how the object looks right after the deletion timestamp is stamped, do the cleanup work, and then remove the finalizer, and only then does the object disappear. There is a delete option that does not wait.

Run several CRs through one pass and build a report

Create two more WebServices, billing and search, in crd-lab (each with an image and replicas), run all three, checkout, billing, and search, through the reconciler for one full pass, and then create /root/op/reconcile/out/reconcile-report.json. items is an array whose elements each have name and action (one of create/update/noop/requeue), summary.noop holds the count of cases where nothing was done, and converged holds true. For all three CRs, status.observedGeneration must equal metadata.generation.

The reconciler is not dedicated to one particular CR. Loop over several targets, collect each judgment, and summarize in one line whether all of them reached the desired state. It has converged only when every target's generation number matches.