GitOps — Give the Repository the Authority to Define State
In one sentence
GitOps is not the name of a deployment tool but a declaration of ownership: the authority to define the state of the cluster lies only in the git repository.
Why this was needed
In a team that deploys with kubectl apply, this question inevitably comes up. "Which commit exactly is running in production right now?" And nobody can answer with confidence. During last week's outage, someone scaled the Pods up with kubectl scale and someone fixed an environment variable with kubectl edit, and those changes were written down nowhere. As time passes, the cluster drifts into a state that is in no repository. This is drift.
The real cost of drift is not "the configuration is slightly different" but non-reproducibility. When you have to rebuild the cluster, applying the repository as it is does not give you the same system as before. In the author's home lab too, when growing from 3 nodes to 7, the newly added nodes carried leftovers of the old cluster (an old kubelet version, old certificates), and mistakes began the moment people started matching those differences from memory.
GitOps flips this problem with a single rule. What is not in the repository must not be in the cluster, and what is in the repository must be in the cluster. Then "what is running now" becomes a question you can answer with git log.
How it works
There are four principles, and you must satisfy all four to be GitOps.
| Principle | Key question | Symptom when violated |
|---|---|---|
| Declarative | Is the desired state written as code? | Non-reproducibility, configuration drift |
| Versioned | Does every change remain as a commit? | No one knows who changed it, when, and why |
| Automatically applied | Is it applied without human hands? | Deployment delay, human error |
| Self-healing | Does it automatically revert when it diverges? | Accumulated drift, environment mismatch |
It is especially important that the third is the pull approach. In the push approach, where CI pushes directly into the cluster, you must put the cluster credentials in CI's hands. In the pull approach, an agent inside the cluster reads the repository, so the credentials never leave the cluster. It is a choice that changes the security boundary entirely.
Two more practical rules are added. First, separate the application code repository from the manifest repository. If you put them in one repository, synchronization is triggered every time the app CI runs, and code review and deployment review get mixed up in one PR. Second, do not use the :latest tag. If the same commit brings up a different image yesterday and today, that repository can no longer define the state. An image tag must be a value that never changes once decided, such as a commit hash or a semantic version.
What you see in the field
First, the commit message is the only basis for a rollback decision. In the middle of an outage, what people look at is not the code but the one screen of git log --oneline. If ten lines of messages like fix and update are piled up, you cannot pick which commit to revert. In GitOps a rollback is a git revert, and choosing what to revert is reading the commit messages.
Second, an urgent hand fix must come back to the code. If you saved the service with kubectl scale, you have to do one of two things: commit that value to the repository, or revert to the repository's standard. If you leave it as it is, the next synchronization quietly reverts it and reproduces the same outage. It is an incident with exactly the same structure as Ansible's -e or a manual change in Terraform.
Third, the repository structure is the permission boundary. The author's home lab keeps the manifests in Gitea (10.0.0.200), and ArgoCD (10.0.0.201) reads them. If you split the directories by app and environment, you can later divide reviewers and deployment permissions along that very boundary. Conversely, if you pack everything into one directory, there is no place to divide permissions.
The three places where GitOps actually gets hard
The principles are simple, but in operation you always get stuck in the same three places.
First, secrets cannot go into git. So the principle "everything is in git" is the first to break. The answer is one of three: upload them encrypted (SOPS, sealed-secrets), refer to an outside store (External Secrets), or inject them from outside the cluster altogether. The first two fit GitOps, and you must start by accepting that either way, there is still a place left to manage the keys.
Second, there are values the cluster changes by itself. The HPA adjusts replicas, webhooks attach annotations, and operators write state. If git reverts those, a never-ending revert loop arises. You write down that such fields are to be ignored.
spec:
ignoreDifferences:
- group: apps
kind: Deployment
jsonPointers: ["/spec/replicas"]
Third, what you fixed by hand in a hurry quietly gets reverted. If you stop a problem with kubectl edit during an outage and forget it, automatic synchronization puts it back a few minutes later. So the habit of leaving in git the fact that you fixed something by hand must be part of the procedure. Turning off automatic synchronization for a while is also a way, but if you forget to turn it back on, from then on git and reality split apart.
A green Synced means "same as git", not "running well". Even if it was applied as the manifest says, a Pod may be in CrashLoopBackOff. Look at health separately, and confirm whether the deployment has really finished by image digest or revision.
Reverting is the biggest benefit of GitOps. Since a revert is just reverting one commit, a record of what changed and how remains by itself. However, changes that cannot be reverted, such as a database schema, are outside this benefit, so release such changes in separate steps that keep backward and forward compatibility.
What you will do in the next lab
You create a manifest repository at /root/gitops/repo, commit it, and apply that declaration to the gitops-lab namespace. Then you deliberately create drift with kubectl scale, see how kubectl diff catches it (exit code 1 if there is a difference), and revert to the repository's standard. Finally, you write with your own hands a synchronization script that does diff, then apply, then records the commit that was applied, and experience exactly what the ArgoCD controller does for you.