The Reconcile Loop — The Power to Revert and the Power to Delete
In one sentence
selfHeal is the power that reverts the cluster toward the repository, and prune is the power that deletes what is not in the repository; the two differ in direction and in risk.
Why this was needed
In the gitops-manifest lab, you made drift by hand. You raised replicas to 5 with kubectl scale, kubectl diff told you with exit code 1, and you typed kubectl apply again to return it to 3. If you ask just one question here, the last piece of GitOps comes out — who types that apply?
If a person types it, that is not GitOps but just a well-organized deployment script. People forget, go on vacation, and get caught up in other outages. What the fourth principle, "self-healing", demands is to make a loop type that apply. A cycle does the job instead of a person's resolve.
How it works
ArgoCD's Application Controller watches resource changes with informers and, for each Application, runs this cycle.
반복:
1. repo-server 에 원하는 상태(렌더된 매니페스트)를 요청
2. 대상 클러스터에서 실제 상태를 조회
3. 정규화한 뒤 둘을 비교 (3-way diff)
4. Sync 상태 갱신 → Synced / OutOfSync
5. Health 상태 갱신 → Healthy / Progressing / Degraded / Missing
6. 자동 동기화가 켜져 있으면 동기화 실행
7. 다음 주기까지 대기 (기본 180초)
Two axes you must distinguish come out here. The Sync status is "is it the same as the repository" and the Health status is "is it running well". It can be Synced and Degraded at the same time — the case where it deployed as the repository ordered but the image could not be pulled and the Pod is in CrashLoopBackOff. Conversely, it can be OutOfSync and Healthy — the case where Pods someone scaled up by hand are running fine. If you read the two axes mixed together, you point at the wrong cause of an outage.
The comparison in step 3 has normalization in it. It strips from both sides the values that Kubernetes fills in automatically (resourceVersion, uid, generation, creationTimestamp, managedFields, and most of status) and then compares. Without this process, a difference would be reported every time even though nothing had changed.
What selfHeal does
selfHeal: true is the switch that moves on to step 6 when OutOfSync comes out in step 4. So a fix made with kubectl scale or kubectl edit disappears at the next reconcile. It is important that it does not disappear immediately — there is a default period, so it feels like "I fixed it, it was fine for a while, and then suddenly it went back". Because of this time lag, it is common to look for the cause in the wrong place.
If you turn selfHeal off, ArgoCD becomes a dashboard that only shows the difference on the screen. In practice, the decision to turn on selfHeal in production is an organizational question of "will we allow urgent manual intervention", not a technical preference. If you turn it on, emergency fixes are quietly reverted, and if you turn it off, drift piles up. Most teams turn it on and, along with it, build the discipline of promoting an emergency fix to a commit right away.
What prune does — and why it is scary
prune: true goes in the opposite direction. The criterion is this.
1. 저장소를 렌더해 "있어야 할 오브젝트" 목록을 만든다
2. 클러스터에서 이 앱이 소유한 오브젝트를 찾는다
(argocd.argoproj.io/tracking-id 어노테이션 = 앱:그룹/종류:네임스페이스/이름)
3. 2에 있는데 1에 없는 것 = 삭제 대상
The scary point is in 3. A single commit that wrongly changes source.path to an empty directory makes the list in 1 zero items, and all the objects that app was managing become deletion targets. It is enough for a reviewer to miss a one-character typo. That is why there are defense devices.
| Device | What it does |
|---|---|
allowEmpty: false |
Refuses synchronization if the render result is empty |
PruneLast=true |
Deletes last, after everything else is reconciled |
Prune=false (resource annotation) |
Excludes only that object from deletion targets |
orphanedResources.warn |
Only warns about objects outside what is managed |
The worst case of selfHeal is "my manual change disappears", and the worst case of prune is "data disappears". The levels of risk differ, so it is safer to attach Prune=false to resources that hold state, such as a PVC.
Fields that must not be reverted
The reconcile loop is diligent, so it reverts even fields that another controller legitimately owns. If the HPA raises spec.replicas to 8, ArgoCD lowers it to 3, and the HPA raises it to 8 again. The solution to this infinite loop is not to turn off a switch but to state the ownership explicitly.
spec:
ignoreDifferences:
- group: apps
kind: Deployment
jsonPointers:
- /spec/replicas
ServerSideApply=true handles the same problem at a different level. It lets the API server track field ownership, so a situation in which someone else manages a field I did not declare is treated as normal.
What you see in the field
First, this lab environment has no ArgoCD controller. The Pod comes up without privileges, and the kwok cluster inside it really runs only etcd, the apiserver, the controller-manager, and the scheduler. So even though you can register the CRDs and create an Application, it does not become Synced by itself. Not hiding this is this curriculum's policy, and so the labs that follow, instead of imitating the controller, have you write by hand in code the judgments the controller makes. You have already done half of the same job — in step 8 of gitops-manifest, the sync.sh you made is exactly a script that goes around steps 3–6 once. The controller is merely something that repeats it without forgetting.
Second, the identity of "why did it go back". This is the inquiry that comes most often in operations. I fixed the configuration, and a few minutes later it is back as it was. The answer is always the same — selfHeal is on, and that change is not in the repository. The one to be angry at here is not the tool but the process. It is because the team agreed to treat a change that did not go through the repository as a change that never existed in the first place.
Third, split the scope of automatic synchronization by environment. A common compromise is to turn on both prune and selfHeal in dev, and in production to turn on only selfHeal or leave it at manual synchronization altogether. This is to prevent the incident where an accidentally merged deletion is immediately reflected in production. Whichever you choose, writing that choice in the Application YAML and committing it to the repository is itself GitOps.
What you will do in the next lab
In the lab right after this, you run this loop around once in eight steps. After declaring the switches and the fields to ignore in an Application, you make, as scripts, the normalization preprocessing, the drift verdict, the deletion candidate calculation, and the empty-result defense. You confirm directly with server-side apply that declared fields come back and fields left to others stay as they are, and finally you create a state where the repository matches and yet the Pod cannot come up, to see why the two axes are separate.