CGOA — GitOps Certified Associate
One Direction Alone Cannot Be the Source of Truth
In one line
The claim that Git is the single source of truth holds only when it works in both directions. If you turn on automated alone, it is one-directional, so Git and the cluster quietly drift apart.
Why it has to be bidirectional
The core claim of GitOps is one thing. Git is the single source of truth, and the cluster is its copy.
For that claim to hold, it must be bidirectional.
| Setting | What it does |
|---|---|
automated |
When you change Git, the cluster follows |
selfHeal |
When you change the cluster, it goes back |
prune |
What you delete from Git is deleted from the cluster too |
With automated alone, it is one-directional. Whatever you fixed directly with kubectl stays, and Git and the cluster quietly drift apart. It is usually because of an urgent incident — you scale up with kubectl scale at dawn and forget to reflect it in Git in the morning. Then at the next deployment, that difference is reverted all at once.
Sync and health say different things
sync— are Git and the cluster the same?health— is what was deployed running well?
It can be OutOfSync yet Healthy. That is when someone fixed something by hand but it runs fine in that state. Conversely, it can also be Synced yet Degraded.
selfHeal has a price
You cannot make emergency fixes. And since what you fixed by hand disappears without a trace, the person who made the fix does not know why their change vanished. That is why the team must agree in advance that "emergency fixes go through Git too."
It also fights with the HPA. If you write replicas in Git, Argo reverts it every time the HPA scales up.
When OutOfSync will not go away
This is the operational problem you meet most often. Git and the cluster look the same, yet it stays OutOfSync. The cause is usually that someone is filling in fields.
| Who fills it in | Example | Response |
|---|---|---|
| Admission webhook | Sidecar injection, default labels | Take that path out with ignoreDifferences |
| Controller | The HPA's replicas, clusterIP |
Remove that field from Git |
| API server defaults | imagePullPolicy, terminationGracePeriodSeconds |
State it explicitly in Git to match |
# Application 스펙
spec:
ignoreDifferences:
- group: apps
kind: Deployment
jsonPointers:
- /spec/replicas # HPA 가 관리한다
- group: ""
kind: Service
jsonPointers:
- /spec/clusterIP # API 서버가 정한다
To see what differs, use commands instead of the screen.
argocd app diff <앱> --local ./manifests # 로컬과 클러스터
argocd app get <앱> -o json | jq '.status.resources[] | select(.status!="Synced")'
Sync order and waves
If you apply all resources at once, ordering problems arise. A CR is applied when its CRD does not exist yet, or the application comes up before the DB is up. Argo decides the order with sync waves.
metadata:
annotations:
argocd.argoproj.io/sync-wave: "-1" # 작을수록 먼저
Argo first applies the default order (Namespace → CRD → the rest) and looks at waves within that. Between waves, it waits until the resources of the earlier wave become Healthy. So if you put a custom resource that has no health assessment in an early wave, it may wait forever — in that case, use a Hook or a health-check Lua script, or merge the waves.
App of apps and projects
When you have dozens of applications, you manage the Application objects themselves with Git (app of apps). Then adding a new service becomes a matter of committing a single manifest.
On top of this you build a fence with an AppProject. It restricts from which repositories, into which cluster, in which namespaces, and what kinds of resources can be created. Without it, a single Application from someone could touch kube-system.
spec:
sourceRepos: ["https://git.internal/labhub/*"]
destinations:
- namespace: "labhub-*"
server: https://kubernetes.default.svc
clusterResourceWhitelist: [] # 클러스터 범위 리소스는 아예 금지
What really matters in practice
Agree on the emergency-fix procedure before turning on selfHeal. Since what you fixed by hand disappears without a trace, the person who made the fix does not know why their change is gone. If you turn it on in a team that has not agreed that "emergency fixes go through Git too," you get a dawn incident twice.
For workloads that use an HPA, remove replicas from Git. If you write it, Argo reverts it every time the HPA scales up, and while load rises, the two controllers fight each other.
Do not read the Synced green light as a completed deployment. Sync means it matches Git and health means it runs well, so a state that is Synced and also Degraded really exists. A deployment verdict must look at both values together.
In the next lab, you stand up a Git server on top of a real Argo CD and trigger these things yourself.