CNPE — Cloud Native Platform Engineer
One green light is not a complete deployment
One-line summary
A completed deployment is not a single status value. It is a conclusion that the intended commit, the applied resources, the ready processes, and the actual request results agree with one another. GitOps faithfully applies even a wrong Git state. So matching Git is different from users being able to use the service.
Why this was needed
A hypothetical platform team changed the readiness path of the order API. The syntax check and the API server's validation passed. The Argo CD screen showed Synced, but the new Pods were not ready, and the ready endpoints disappeared from the Service. If you guess at this point, "it matches Git, so it must be a network problem," you miss the mistake you just deployed. Conversely, if you roll back first for every outage, you can make a data change or an external dependency problem more complicated. First, gather evidence of what changed.
A GitOps controller is not a referee who knows the business answer. It is a controller that applies and observes the declared state. The health judgment of an application also depends on per-resource-kind rules. Even when Healthy is displayed, an API that returns a wrong price or fails only for certain users can remain. A simple HTTP check in a learning environment likewise does not prove every feature of a real service.
How it works
Read the four layers together
| Question | Evidence to check | What this alone cannot tell you |
|---|---|---|
| What did we intend to deploy | The Application's targetRevision, the Git commit and its change history | Whether it was actually applied |
| Does it match that commit | status.sync.revision and Synced | Pod readiness and request success |
| Does the controller consider it ready | health, the Deployment generation, replicas, and conditions | Correctness of the business response |
| Do requests actually work | The HTTP status via the Service and the expected body | Success for every path, user, and load |
For a single-replica Deployment used in the lab, start by checking whether the desired replicas is 1. If the current generation differs from observedGeneration, the controller may not yet have observed the new declaration. Look at updated, ready, and available together, and check whether the Service's EndpointSlice has a ready endpoint. If you conveniently turn an empty JSON or a missing field into 0 or success, you can hide an outage. When the needed information is missing, you should defer the judgment or treat it as a failure, and show why it is insufficient.
The following is an example of diagnostic commands in an isolated lab environment you built yourself. Replace demo with the real
Application name and namespace. These are not commands to deliberately break a production environment.
kubectl -n argocd get application demo -o json
kubectl -n demo get deployment web -o json
kubectl -n demo get endpointslices -l kubernetes.io/service-name=web -o json
kubectl -n demo describe pod -l app=web
After reading the latest conditions and events from this output, send an actual request. A request sent inside the VM to the ClusterIP
verifies only that path. External users may go through DNS, a gateway, TLS, and authentication,
so in production you also need observation along the same path as users. Do not restate a request sent with kubectl exec to the Pod
itself, which succeeded, as success of the Service or the external path.
The same Synced gives opposite results
Suppose the readiness of a healthy nginx checks /. You change it to a nonexistent
/not-ready and commit it to Git. The YAML is valid and the container process runs,
but the HTTP probe fails. Even when the new declaration is applied and becomes Synced, the Pod is not Ready.
When the progress deadline is exceeded, you can check the Deployment's Progressing condition and its reason.
In a dedicated reproduction environment, you can use 1 replica and the Recreate strategy to see the difference clearly. This is not a recommended availability setting but a deliberate condition for observing the failure. With RollingUpdate, if the old Pod keeps serving, the HTTP response can still be 200 even if the new Pod fails. So "deploying a bad probe always kills the whole service" is also a wrong generalization. You must look at the strategy and the remaining healthy Pods together.
Commit pinning and recovery
main tracks the latest commit of the branch. An Application pinned to a specific commit SHA
does not automatically move to a new commit just because you pushed it to the branch. In this case you must also update
targetRevision to the intended new SHA. Even if you pin the Git SHA, if the image in the manifest is
a moving tag, you have not pinned the image content. Distinguish what each one identifies.
One way to recover from a stateless configuration mistake is to create a new commit that cancels the bad change with git revert
and deploy it. The content of the recovery commit can be the same as the earlier healthy commit, but the SHA is
new. So the judgment "recovery succeeds only if it equals the original SHA" is not appropriate.
Check together the new SHA you chose as the recovery target, that SHA's tree, and the cluster's observed state and responses.
In an Application with selfHeal turned on, an emergency fix that edits only the Deployment directly can be reverted to the wrong state in Git. For this reason you must fix the intent in Git as well. The Argo CD rollback command in an auto-sync state and reverting in Git and then deploying a new commit are not the same operation. The former is a rollback action the tool provides, and the latter is a way of fixing the source of the declaration with new history.
What it looks like in the field
For the hypothetical outage above, the recovery report records the healthy, failed, and recovery commits separately. At each point in time, keep the sync revision, health status, probe failure reason, endpoints, and HTTP result together. If only a single Healthy screenshot is left after the outage, it is hard to connect the cause of the outage to the change that recovered it. Do not put sensitive request bodies or tokens into the evidence files.
Test the promotion check in both directions too. Do not feed it only responses from the healthy state; feed it a 200 from the previous commit, a state that is Synced but Degraded, a state with no endpoints, a 200 with the body of a different release, and a state with a missing required field, and confirm that each is rejected. This is separate from a test that runs a real controller. The former looks at the judgment logic, and the latter looks at whether that state actually occurs in a real environment.
This example is the recovery of a stateless web configuration. It is not a procedure for rolling back a DB schema or canceling messages that have already been sent. A release with data changes needs separate design for backward compatibility, backup and restore, reprocessing, and duplicate handling. Do not claim data recovery on the basis of a green light on a screen.
What to do in the next lab
In the existing CNPE delivery gate lab, read the file and declaration checks and the actual-execution checks as distinct. Do not judge that a real Argo CD sync was performed merely because you wrote an Application file. In the next VM lab, you observe normal, failure, and recovery directly in a real Argo CD, and leave Git history and persistent evidence files. At the end, you build an approval code that rechecks even the current service. Do not extend the evidence of stateless web recovery to DB recovery.