KCNA — Kubernetes and Cloud Native Associate
What readiness failures and manual rollbacks leave behind
One-line summary
A progress failure condition is not an automatic recovery command, and a Deployment rollback is not a transaction that also rolls back external configuration and data.
Why this was needed
The new v3 app has a live process and an open port. But its readiness and business path return 503. If you had left only the new version and deleted the old Pods first, users would have received errors. This experiment keeps two healthy v2 Pods and adds only one v3 that is not ready. It is a scene where "scaling up the new version quickly" and "protecting current users' requests" conflict.
If the Service keeps delivering v2, the operator can catch their breath for a moment. But you must not assume that just waiting will certainly fix it. Deleting the probe to make a green light while the probe is correctly failing is not a recovery either. What is needed is to decide which version and configuration to return to, and to check that actual requests have returned to that state.
How it works
This Deployment uses 2 replicas, maxSurge 1, and maxUnavailable 0. During a rolling update it can add one more than the desired number, but it gives the update no room to arbitrarily reduce existing available replicas. If the new Pod is not ready, the controller can keep the two still-healthy old Pods. So a state is possible where desired is 2 but there are actually 3 Pods and available is 2. This is an experiment with integer settings, and we do not generalize the rounding rules of percentages from the numbers here.
Readiness judges readiness to receive requests. The liveness of this synthetic app stays healthy, so v3's readiness failure alone does not make the container restart repeatedly. If you read the Pod UID, container ID, restart count, and the direct HTTP response together, you can confirm "running but not ready." The Service should deliver requests to the ready v2 and exclude v3. Even maxUnavailable 0 is not a zero-downtime guarantee that blocks every external failure, node failure, and wrong probe.
progressDeadlineSeconds is the budget for deciding that a deployment cannot make progress. This experiment uses 20 seconds to shorten the observation time, but the appropriate value in production has to be set according to the real budget of image download, app start, and minimum readiness time. After the deadline passes, you can observe a state where the Progressing condition is False with reason ProgressDeadlineExceeded. This does not mean the new version has turned healthy, nor that Kubernetes has rolled back to the old template on its own.
Even after the failure condition appears, you check whether spec.template is still v3, whether the healthy v2 UIDs and containers are unchanged, and whether the new v3 is still not ready. Investigating the last healthy revision and rolling back manually is the decision that comes after. In the lab you choose the v2 revision you already observed, and you do not memorize simply the smallest or the most recent number.
kubectl -n kcna-delivery rollout history deployment/web
kubectl -n kcna-delivery describe deployment web
# 조사한 정상 번호를 사용합니다. 아래 N은 그대로 실행하는 값이 아닙니다.
kubectl -n kcna-delivery rollout undo deployment/web --to-revision=N
A rollback is a new change that goes back to a previous template. Do not expect the history number to simply decrease to a past number. Also, in this experiment the remaining healthy v2 ReplicaSet and Pods can be reused. Do not generalize that every rollback must create new Pods, or that it always keeps the existing Pods. The actual transition differs depending on which ReplicaSets remain and how many replicas are alive.
What it looks like in practice
One reason an incident does not end even though the code rollback succeeded is external state. In this experiment, before deploying v3 you change the ConfigMap to purple. The existing v2 process preserves the environment variable green, but the configuration object itself is purple. Even if you roll the Deployment back to v2, the ConfigMap stays purple. If you let it pass because the response is green for now, a replacement Pod created later may receive purple. That is why, after a rollback, you also investigate the declared configuration and the running configuration.
Database schemas, external service configuration, and messages already sent are larger boundaries. The Deployment history does not atomically rewind them. For a rolling update in which old code and a new schema must live together, compatibility design has to come first. This is why an approach of deploying in stages is needed: adding a new column, moving the code, and removing unneeded columns. This experiment does not prove the safety of database migrations either.
In a production environment managed by GitOps, the desired recovery state must also be reflected in Git's declaration. If you only make a manual change and leave Git as it is, the reconciler may apply it again. Since ArgoCD is not installed on this personal k3s, we do not claim to have seen self-heal actually revert it. First you learn the Deployment's behavioral boundaries separately, and in the GitOps lab you add the source of the desired state.
What you will do in the next lab
After it has been changed to a healthy v2, you put in a v3 that fails readiness and record that it was not rolled back automatically even after the deadline was exceeded. After recovering manually to the healthy revision, you check whether the ConfigMap remained purple and restore it to green separately. Without deleting policies or touching the healthy comparison group, you compare three pieces of evidence: the actual response, the object identities, and the configuration.
Official sources: Deployment rollback and progress conditions · How ConfigMaps are consumed.