TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Recovery must survive the next pod

Continue in TT Lab

In one sentence

Safe recovery is not proven by a single healthy response. The controller status that observed the current declaration, the Pods that came from that declaration, the actual response, and the replacement Pod's configuration must all line up.

Why this was needed

Even when a new Pod fails its readiness check, the service can be healthy. That is because the earlier Pods keep handling requests. This is not necessarily an abnormal state; it can be the result of design so that a bad deployment does not remove the existing capacity all at once.

The problem is when you see a healthy HTTP response and conclude that the new version succeeded too. If an existing Pod responded, that is not evidence of the new version's success. Conversely, you must not conclude that every user is suffering an outage just because one new Pod is not ready. Deployment progress and user impact are connected, but they are not the same value.

In the actual probe, two existing healthy Pods and one new Pod that was not ready existed together. The service responded normally from the existing Pods. The first rollback completed, but after the existing Pods were replaced, some capacity was again not ready. This is where the grounds came from for including Pod replacement in the definition of recovery.

How it works

First confirm that the declaration was read

An API response that saved a change does not mean the controller has already processed it. Read the Deployment's metadata.generation and status.observedGeneration together to see whether the current declaration was observed. Next, distinguish the new template's replica count, ready count, and available count. Only the Available condition from the previous state may remain.

This example uses two replicas, maxSurge: 1, and maxUnavailable: 0. It is a setting that makes sure the existing healthy capacity is not reduced first when the new Pod is not ready. A terminating Pod may appear additionally for a moment, so do not confuse the length of a simple list with the deployment's capacity calculation.

You must also express precisely the fact that an observation command did not finish within the set time. It means you did not see completion in that observation window. It is separate evidence from a progress-deadline-exceeded report by the deployment controller or from a readiness failure of the actual app. If you restart the same installation based only on a timeout, you may erase the cause or overlap runs.

Connect who the responding Pod is

The example app includes the Pod name, the configuration version it read, and the health status in its response. Trusting only these values is not enough. You also cross-check by UID that the Pod you queried belongs to the relevant ReplicaSet and that ReplicaSet belongs to this Deployment. Names can be reused.

Look at each individual Pod's /readyz response together with Kubernetes's Ready condition. Liveness, meaning the process is alive, differs from readiness, meaning it may receive traffic. In the example, even with the bad configuration the process is alive, but the readiness response is 503. You also connect which healthy Pod the service's successful response came from.

This does not mean a real-world app must always expose such information to external users. This small educational response is a device to reveal the relationships among the observed objects. In production, you choose observation means with controlled access, such as internal diagnostic paths, logs, and deployment metadata.

Test recovery with the next Pod

After returning to the earlier version, replace one healthy Pod. After confirming that the old UID is gone and a new UID has appeared, check whether that new Pod actually read the earlier configuration and became ready. Do not reuse the earlier healthy Pod's response as the new Pod's success.

This replacement is done only on the designated workload inside your own personal VM. It is not a procedure of deleting an arbitrary Pod in production to confirm recovery. In production, you must first consider capacity, the disruption budget, traffic, and the approved test scope. The probe's delete request also carried UID and resourceVersion preconditions so that it would not ignore other objects or concurrent changes.

What it looks like in the field

A platform team must provide clues for understanding failure at the same time as it provides deployment tools. Rather than the single line "deployment failed," an explanation like "one new replica is currently running with configuration v2 but its readiness response is 503, and the two existing v1 replicas are serving" is far more useful for developers to decide their next action.

In a recovery record, leave not only what was changed but also what was preserved. This experiment re-confirmed at the end the UID and healthy response of another team's Deployment. This does not prove all tenant isolation, but it rules out a false success in which another app is deleted to revive my own.

When saving records, link the current run to the target UID, and make sure later steps do not overwrite earlier observations. Do not record a lookup failure as an empty list or a healthy state. Only if the meaning of earlier evidence is preserved even when the same step is graded again after completion can learners look back at what they learned.

What to do in the next lab

You perform a deployment and rollback with mutable configuration and observe the recurrence after Pod replacement. Next, you recover using per-version immutable configuration references and compare the result in which even the replacement Pod is healthy. The goal is to read the command's exit code, the controller status, the actual response, and the configuration source as different pieces of evidence.

Official documentation