CGOA — GitOps Certified Associate
refresh Is Computation; sync Is Application
In one line
A refresh only re-renders and computes the diff; it applies nothing. Only a sync touches the cluster. If you miss this one sentence, half of Argo CD will not make sense.
Why this was needed
When you press the Refresh button in the UI and the app is still OutOfSync, most people think, "Is it broken?" It is normal. Here is what a refresh does.
- It asks the repo-server to regenerate the latest manifests (depending on whether the cache is invalidated; with a hard refresh, it starts again from the clone).
- It re-reads the live state from the target cluster.
- It compares the two and updates the Sync Status (Synced / OutOfSync) and the Health Status.
That is where it ends. Applying is the job of a sync, and a sync happens only when a person presses it or when syncPolicy.automated is turned on. Thanks to this separation, you can keep observing "how different are Git and the cluster right now" without changing anything. This is exactly why many organizations operate with automated sync turned off — always watch, but let a person do the applying.
How it works
3-way diff — why three states?
Argo CD's diff looks at three states, not two.
| State | Where it comes from | What it answers |
|---|---|---|
| Desired | The result of rendering Git | What we want |
| Live | The object currently in the cluster | What is real right now |
| Last-applied | The last-applied-configuration annotation of the object |
What we declared previously that we would manage |
Comparing only two states leads to a fatal misjudgment. Suppose, for example, that an HPA has raised replicas to 5 and Git does not have the replicas field at all. If you compare only Desired and Live, the conclusion is "it is a field only in Live, so I should delete it." Looking at the third state changes the answer — replicas is not in last-applied either, so it is a field we never managed in the first place, and since it is a value owned by someone else's controller, we must not touch it.
Normalization comes into play here too. metadata.resourceVersion, uid, generation, creationTimestamp, managedFields, and most of status are removed from the diff. Defaults that Kubernetes fills in automatically — a Service's clusterIP, and imagePullPolicy: Always when the image tag is latest — are also ignored. Without this normalization, every app would look OutOfSync forever.
What selfHeal really means
syncPolicy.automated.selfHeal: true means "automatically revert drift." Flip it around and it means this.
Anything you fixed by hand during incident response comes back.
If a Pod died at dawn and you urgently raised replicas with kubectl scale, it goes back to the Git value on the next reconcile loop. This is not a bug; it is the design. An organization that turns on selfHeal has also accepted the discipline that "emergency fixes are made by commit too." If you need an escape hatch for urgent moments, the standard approach is to temporarily disable automated sync, or to put the field in ignoreDifferences so it is taken out of management.
Waves and hooks — two mechanisms that create ordering
Declarative cannot express order. So two mechanisms are layered on top.
A sync wave groups resources by the number in the argocd.argoproj.io/sync-wave annotation and applies them from the lowest number. The key point is that it does not move on to the next wave until all resources in the current wave are Healthy. That is why dividing waves badly stops the deployment right there — for example, if you put a resource that never becomes Healthy (such as a Job that is never ready because it gets no incoming traffic) in an early wave, the later ones never arrive. Within the same wave, the default order by resource kind applies (Namespace → NetworkPolicy → ResourceQuota → LimitRange → ServiceAccount → Secret/ConfigMap → RBAC → CRD → PV/PVC → Service → workloads → Ingress).
A hook divides the phases themselves. The order is PreSync → Sync → PostSync, and if it fails, SyncFail runs. A hook is usually a Job, specified with the annotation argocd.argoproj.io/hook: PreSync. The deletion policy argocd.argoproj.io/hook-delete-policy has three values, HookSucceeded, HookFailed, and BeforeHookCreation, and the default is BeforeHookCreation. That is, a hook resource stays around even after it succeeds and is deleted just before it is created anew at the next sync. This default is the reason you can look at the logs of a failed migration Job after the fact.
Retries and backoff
When a sync fails, it retries with exponential backoff. With duration: 5s, factor: 2, and maxDuration: 3m, the wait grows as 5s → 10s → 20s → 40s → 80s and hits the cap at 3 minutes. The situations that trigger a retry are a resource apply failure, a Health Check timeout, a hook Job failure, and a transient network error.
The danger of prune
prune: true also deletes from the cluster the resources that have disappeared from Git. The criterion is not "everything that is in the cluster but not in Git," but what is not in Git among the things Argo CD has marked as its own. The marking mechanism is resource tracking, and the default is the annotation method.
argocd.argoproj.io/tracking-id: APP_NAME:GROUP/KIND:NAMESPACE/NAME
예) checkout-prod:apps/Deployment:cgoa-prod/prod-checkout
The legacy method uses the label app.kubernetes.io/instance. Because other tools such as Helm also use this label, ownership decisions can collide, so the annotation method is recommended.
The reason prune is frightening is that a single commit that changes a path wrongly is itself a mass deletion. If you mistype source.path so that it points to an empty directory, the render result becomes 0 resources, and every resource that app managed becomes a prune target. The defenses are allowEmpty: false (reject an empty render result), PruneLast=true (prune last, after the other resources have finished syncing), and, on individual resources, argocd.argoproj.io/sync-options: Prune=false.
What it looks like in the field
In the author's homelab, the moment this sense was needed was when expanding the GPU nodes. GPU Feature Discovery automatically attaches labels to new nodes — an RTX 3090 24GB, a 5090 32GB, and two 4070 Laptop 8GB. But if a Pod requests only nvidia.com/gpu: 1, a training job that needs 32GB can land on an 8GB laptop GPU. To Kubernetes, both are "one GPU."
So semantic labels were added by hand (along the lines of gpu.homelab/tier=xlarge and vram=32g). Here comes the lesson from the GitOps point of view — labels attached by a controller and labels declared by a person coexist on the same object. If you manage this node object with GitOps, the labels GFD attached must be taken out with ignoreDifferences. If you do not, the reconcile loop and the controller fight, erasing each other's fields. This scene holds everything about why the 3-way diff and field ownership are needed.
What you will do in the next lab
After putting a manifest on a real cluster, you change replicas by hand to create drift, leave that difference in a file with kubectl diff, and then revert it (a person doing self-heal). Next you write, as files, three manifests with sync wave annotations and a PreSync hook Job, and finally you search the cluster and determine for yourself what the prune targets are.