TT Lab
Get started
Learn Learning paths Courses

CRDs and Operators

Reconcile — Observe, Compare, Act, Report

Continue in TT Lab

In one sentence

Reconcile does not ask "what happened." It asks only "what is the desired state now, and what is the actual state," and closes the difference idempotently.

Why it was needed

Imperative automation scripts are a list of "create this, configure that, then run this." If one fails midway, the system is left in a half-finished state, and running it again produces an "already exists" error or creates duplicates. So people come to fear rerunning, and automation people fear ends up unused.

Reconcile turns this around. Wherever it fails, the next call runs again from the start and finishes only the part that was not done. It does not matter if the controller dies midway and comes back. There is no need to hold "how far I got" in memory, because the real state is always in the cluster.

How it works

In a real controller, a request reaches the reconciler along three stages.

API Server --watch--> Informer --> WorkQueue --> Reconciler
                       (로컬 캐시)  (중복 제거,   (사용자 코드)
                                    속도 제어)

The most important insight here is that what is passed to reconcile is not the object but only the object's key (namespace/name). What changed, and whether it was a create or an update, is not passed. Reconcile re-reads the latest state using that key. This design enforces idempotency.

A single reconcile boils down to six questions.

  1. What is the desired state of this object now (desired)
  2. What is the actual state of the cluster (observed)
  3. What is the difference between them (diff)
  4. To close that difference idempotently (action)
  5. Is there anything not yet closed (requeue decision)
  6. How to reflect the observed result in status (status)

The idempotency of step 4. It must be written not as "create" but as "check that it exists in the desired form, and if not, make it match." If it already matches, doing nothing is the right answer, and the resourceVersion of the child object must stay the same. If you write again with identical content, a new watch event is generated, and that event calls reconcile again.

The requeue of step 5. You must not treat every "not ready yet" as an error. If you return it as an error, the work queue retries with exponential backoff and error logs pile up, and that object's backoff grows longer, which reduces responsiveness. A dependency that does not exist yet is not a failure; it means "let's look again in a moment."

Kind Trigger Wait time Purpose
Error backoff Returning an error Grows exponentially (automatic) Retrying real failures
Delayed requeue Stated in the result The value I specify Intentional periodic checks

The reason backoff is exponential is so that repeated failures of one object do not starve the whole queue. And when it succeeds, that key's backoff counter is reset.

The status of step 6 and infinite loops. Writing status can itself become a new watch event and call reconcile again. Real controllers prevent this with GenerationChangedPredicate. Because generation rises only when spec changes, applying this filter screens out the re-invocations caused by the controller's own status writes. But if the controller needs periodic checks, take care that this filter does not block those checks too.

The deletion path and finalizers. Children inside the cluster linked by owner references are cleaned up by the garbage collector, but resources outside the cluster (a cloud load balancer, an external DB account) are unknown to Kubernetes. A finalizer is the mechanism that guarantees that cleanup.

삭제 요청 -> deletionTimestamp 설정 (오브젝트는 아직 존재)
          -> 리컨사일 재호출: 정리 수행
          -> 파이널라이저 제거
          -> 그제서야 실제 삭제

The key point is that "delete" does not mean "vanish immediately" but "delete = a deletion timestamp is stamped and reconcile is called one more time." That one time is the last chance to clean up. And the cleanup function must also be idempotent — if an attempt to delete something that is already gone throws an error, the finalizer never comes off and a deletion deadlock results.

An honest note about this lab environment

This lab has no way to run a compiled Go controller. So we build the reconcile loop as a shell script reconciler. We do not simulate the informer cache or the work queue; instead, we implement by hand the essence of the reconcile function: observe, compare, act, then status reporting and the requeue decision, plus the finalizer flow. What is graded is not a controller binary but the correctness of the judgment and idempotency. Even if you port a real controller to Go, these four beats stay the same.

What you see in the field

First, infinite reconcile. It is a rite of passage that almost every Operator developer goes through once. When reconcile writes status, that write creates a new watch event, and that event calls reconcile again. The CPU pins a core and logs pile up by the hundreds of lines per second. There are two fixes — either attach a predicate that looks at metadata.generation, which rises only when spec changes, to cut the self-triggering caused by status changes, or write status only when the value has actually changed in the first place.

Second, status conflicts. You try to update status with a stale object read from the cache and hit a conflict. The probability rises when the controller handles several resources concurrently. Re-read the latest object before updating, or leave it to server-side apply for field-level merging.

Third, the object you just created doesn't show up. Reads go to the informer cache and writes go to the API server. So if you Get right after a Create, you may still get "not found." Adding a sleep here is the most common wrong answer. The right answer is not to wait — write it idempotently and let the next reconcile see it naturally. This is where the design philosophy of the reconcile loop shows itself directly.

Fourth, an object stuck on a finalizer. If you delete a CR while the controller is dead, it stays in Terminating forever with only deletionTimestamp stamped, because nobody is there to remove the finalizer. That is why, when removing an Operator, you must follow the order of cleaning up CRs before the controller, and if you are in a hurry, write the escape hatch of removing the finalizer by hand into the runbook.

What you will do in the next lab

You write /root/op/reconcile/reconcile.sh, which reads the CR's spec as desired and the child ConfigMap as actual, compares them, creates it if missing, and proves that it does nothing when it already matches, using the fact that resourceVersion stays the same as evidence. You write status back through the subresource path, leave the requeue decision for when a dependency is missing as JSON, build a flow where an object is cleaned up and then deleted through a finalizer, and finally produce a reconciliation report from running three CRs through one full pass.