CGOA — GitOps Certified Associate
Four Green Lights, Four Different Questions
In one line
Matching Git, Kubernetes being ready, and a user's request succeeding are three different claims.
Why this was needed
A message arrives in the on-call channel: "Payments aren't working." The deployment screen shows Synced and Healthy, and the Pods are Ready. If you answer here, "The infrastructure is fine, so the user should just retry," you have drawn a conclusion larger than the range you checked. Nor do you need to assume that the green on the screen is wrong. The question that indicator answers and the question the user is asking can be different. This lab does not connect a payment system. It reproduces this difference with a small model in which nginx's business path deliberately returns 500. It uses the Argo CD and HTTP server of a personal VM, not a fake status file made to look like a real incident response.
How it works
When investigating, separate the following four questions and take them in order.
| Question | Evidence to check | What this alone does not tell you |
|---|---|---|
| Which change did we try to deploy? | The files in the remote main and the Git commit | Whether the controller saw that commit |
| Has that target been reflected on the cluster? | The Application's sync revision and Synced | Whether the running process read the new configuration |
| Does the resource meet its readiness conditions? | health, observedGeneration, updatedReplicas, Pod Ready | Whether the business path returns the correct body |
| Does the request we are investigating succeed? | Request path, response code, body, observation time | The success rate across all external users and long-term availability |
Argo CD's default health assessment uses the Kubernetes status for each resource kind. For example, checking that a Deployment's current generation has been observed and its updated replicas is important, but it does not automatically know the order semantics of a shopping mall. Synced is also not a feature that evaluates the meaning of an application. If you stored a wrong configuration in Git, the result of faithfully deploying that configuration can be Synced. So GitOps's declarative management and good application verification do not replace each other. Official resource health documentation.
Readiness is a signal for judging whether this Pod is ready to receive traffic. But what it actually checks depends on the probe we wrote and the app's implementation. In this lab, /healthz returns only a 200 that means the process responds. The business path / returns 500. This design is deliberately imperfect and is not a recommended production probe. That said, changing things so that every Pod's liveness fails when an external payment API is briefly slow does not fix anything either. Readiness's traffic-delivery control and liveness's restart decision have different purposes. If you link the two signals without checking whether an external dependency failure can be fixed by a restart, you can create restart load instead of recovery. Official probe concepts.
What it looks like in the field
The first trap is looking only at the green screen without checking the commit. You may be looking at the Healthy of a previous commit rather than the change you just pushed. Compare the local HEAD, the server's remote main, and the Application's status.sync.revision together. The branch name main is a moving name, and the commit SHA is the concrete source at that moment. The fact that the push command finished is not evidence that the deployment is complete. In this lab you ask for a re-comparison of the relevant Application with a refresh, but that too is not a declaration of success, so you wait for the actual revision and the state to converge. You do not change the overall controller settings or other Applications.
The second trap is leaving only the response code. Even a 200 may be a stale body or a response from a different version. If you leave the service address, path, body, and time together and link them to the deployed Pod's UID, the claim becomes clearer. Conversely, a 500 is an observation that the server sent an HTTP response. Do not lump a DNS failure or a connection failure into the same event. The 500 in this lab is reproduced with the nginx configuration, so writing an external API or DNS failure as the cause would not match the observation.
The lab calls a ClusterIP from the VM. The success of this path alone does not let you say that external DNS, Ingress, TLS, and authentication are healthy. In the final report, write the conclusion that the external-path check remains. Not turning the result of three successes in a short time into a month of availability or an actual payment success rate is also part of an engineer's ability to verify.
What you will do in the next lab
In the first observation, you preserve the Git revision and the resource identities, and read readiness and the business response separately. You narrow down the cause from the exposed nginx configuration and the actual response. After that, you fix the configuration in Git, but whether fixing the configuration alone changed the existing process too is confirmed by a separate observation. Do not make up numbers or UIDs; investigate from the actual observations the helper left.