Why You Need Three Things to Compare
In one line
Argo CD's Application controller wakes up every 180 seconds by default and compares three states: Desired (Git), Live (the cluster), and Last-Applied (the annotation). The whole of this lesson is why there are three rather than two, and why, without normalization before the comparison, it stays OutOfSync forever.
Why this was needed
Why not just compare Git and the cluster? Because you cannot. When Kubernetes stores an object, it fills in a large number of defaults. If you submit a Service without a clusterIP, the API server assigns one; if the image tag is latest, it puts Always in imagePullPolicy; and it attaches resourceVersion, uid, generation, creationTimestamp, and managedFields to every object and fills in status. The Git manifest has none of these. So the conclusion "they differ" always comes out, and if you leave auto-sync on, the controller repeats a meaningless apply forever every 3 minutes.
The reason a third comparison target is needed is a different kind of problem. If you know only Git and the cluster, you cannot tell whether a "field that is not in Git but is in the cluster" is something I put in and later removed, or something another controller (an HPA, a sidecar injector, a defaulter) put in. Last-Applied remembers "what I declared last time." A field I declared before that has now disappeared from Git should be deleted, and a field I have never declared belongs to someone else and must not be touched. The 3-way diff makes this judgment. For this calculation, Argo CD uses the same Structured Merge Diff library that Kubernetes Server-Side Apply uses.
How it works
The representative fields removed in the normalization step are metadata.resourceVersion, metadata.uid, metadata.generation, metadata.creationTimestamp, metadata.managedFields, and the status of most resources. Differences that still remain must be excluded by hand with ignoreDifferences. The most common case is the HPA. If Git says replicas: 2 and the HPA has raised it to 8, it is OutOfSync forever, and if selfHeal is also on, a fight begins in which Argo CD lowers it to 2 and the HPA raises it to 8. The answer is to exclude /spec/replicas from the diff.
Sync is divided into three phases: PreSync → Sync → PostSync, and if Sync fails, a separate SyncFail phase runs. What runs in each phase is a hook, and the default deletion policy of a hook resource is BeforeHookCreation. This means that on the next sync the previous hook object is deleted first and then created anew, so the traces of a failed migration Job remain until the next deployment.
The order within the same phase is decided by waves. They run from the lowest number, wait until all resources of a wave become Healthy, and then move on to the next wave. If any wave fails, the whole sync stops. This "wait" is important. A wave is not a device for changing apply order but a device for waiting for readiness.
Retries use exponential backoff. With limit 5, duration 5s, factor 2, and maxDuration 3m, it tries five times at intervals of 5, 10, 20, 40, and 80 seconds. And the way Argo CD recognizes its own resources is the tracking annotation, whose format is APP_NAME:GROUP/KIND:NAMESPACE/NAME. The core group has an empty group name, so the slash comes right after the colon, as in my-app:/Service:default/nginx-svc. If you write this format by hand, you get a feel for why the annotation method is recommended over the label method. A label has a 63-character limit and does not include the namespace.
What it looks like in the field
The author's homelab ArgoCD runs at 10.0.0.201 obtained from the MetalLB pool, and Gitea is at 10.0.0.200 in the same range. There is one trap you must point out when registering a cluster: argocd cluster add creates an argocd-manager service account in the target cluster and by default binds it to cluster-admin. You can let it pass in a homelab, but if you leave it like this in production, one GitOps controller becomes root on every cluster. The standard approach is to create a separate ClusterRole containing only the needed apiGroups and verbs and swap the binding.
Another thing: this cluster has Cilium Gateway API running, and when the Gateway API CRDs were kept at v1.2, the controller refused to start, saying tlsroutes and referencegrants were not v1. It came up only after upgrading to v1.6.1. This kind of thing happens often when deploying CRDs with Argo CD: if a CRD and the custom resources that use it are in the same sync, the CR is applied before the CRD is registered and fails. Splitting them into waves and sending the CRD first is the standard solution, and it is the best example of why waves exist.
What you will do in the next lab
You build up an Application manifest in /root/capa-app/ one field at a time. You fill in source, destination, syncPolicy, retry, and ignoreDifferences in turn, and at the end you actually put the namespace and Deployment that the app will deploy into the cluster and attach the tracking annotation by hand.