TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

Why does the next pod fail after a successful rollback?

Continue in TT Lab

Goal

On a real app with the same image, you compare configuration changes, rollouts, rollbacks, and Pod replacement. You observe the recurrence of a rollback that references mutable configuration, and confirm whether recovery that references per-version immutable configuration holds even on a new Pod.

Why it matters

A successful rollback command and a healthy service response alone cannot tell you that the external configuration was restored as well. While the old Pods respond normally, a new Pod can receive a bad configuration. This is a personal VM that uses a real k3s, kubelet, and containers, not KWOK or a fake Ready state. Installation can take several minutes. It is a 55-minute lab, and if you need more, extend the time before it expires. The VM and files are reclaimed when the session ends. Download any records you need first. This is not a procedure for replacing Pods in a production cluster.

Prepared environment and tools

act performs only the change you specify. capture observes for up to 75 seconds and then saves JSON, and observe queries once. grade does not modify student files or resources. Completed past observations are not overwritten with the current state of later steps. The record JSON is not an answer you dress up yourself but data saved after a successful observation. Write only the input JSON with the specified values. The release JSON in steps 3, 6, and 7 is a strategic merge patch rather than a whole Deployment, and it keeps the remaining security settings of the container with name=app.

Steps

  1. Save /root/cnpa-release/baseline.json with capture 1. Check that the two actual Pods of the cnpa-rollback/checkout Deployment read configuration v1 and are ready, and that the Pod in the service response is owned by that Deployment.
  2. Write the v1 ConfigMap settings (namespace cnpa-rollback) in /root/cnpa-release/mutable-config.json. Its data is CONFIG_REVISION="v2", HEALTHY="false", and immutable=false. Apply it with act 2 and save env-unchanged.json with capture 2. The configuration has changed, but the UIDs of the two initial Pods and their actual v1 environment variables must be preserved.
  3. In /root/cnpa-release/mutable-release.json, write a JSON patch that sets only labhub.io/release in spec.template.metadata.annotations to mutable-v2. Save rollout-incomplete.json with act 3 and capture 3. One new Pod must be v2 and not ready (503), the two existing ones must be v1 and healthy, and the service responds from the existing Pods.
  4. Roll back with act 4 to the previous revision observed at installation, and save apparent-rollback.json with capture 4. Check together that rollout status succeeds and the first two Pods are healthy again, and that the contents of settings are still the bad value.
  5. Replace one healthy Pod in your own lab with act 5 and save replacement-failed.json with capture 5. Check that the deleted UID is gone and that the Pod with the new UID receives the bad v2 configuration. One existing Pod must still respond normally.
  6. In /root/cnpa-release/versioned-configs.json, write two ConfigMaps as a v1 List. In the cnpa-rollback namespace, settings-v1 is CONFIG_REVISION="v1" and HEALTHY="true", settings-v2 is CONFIG_REVISION="v2" and HEALTHY="false", and both are immutable=true. In versioned-release.json, write a patch that specifies the template annotation labhub.io/release=versioned-v1 and envFrom.configMapRef.name=settings-v1 for the container with name=app. Test applying it and the rejection of an immutable modification with act 6, and save versioned-ready.json with capture 6.
  7. In /root/cnpa-release/bad-release.json, write a patch that specifies the template annotation labhub.io/release=versioned-v2 and envFrom.configMapRef.name=settings-v2 for the container with name=app. Save versioned-incomplete.json with act 7 and capture 7. The two existing immutable v1 Pods must be healthy, and the one new v2 Pod must fail to become ready.
  8. With act 8, recover to the actual revision from step 6 and then replace one healthy Pod. Save replacement-ready.json with capture 8. Check that the replaced new UID is also healthy with the immutable v1 configuration, and that the UID of settings-v1 and the UID and HTTP response of the other team's cnpa-rollback-other/sentinel are preserved.

Notes

A Deployment revision is the history of the Pod template, not a copy of the external configuration contents. Do not guess the number; use the revision you observed. Connect the service response, the individual Pod responses, the current observed generation, and the UIDs of Pod→ReplicaSet→Deployment. A terminating Pod may appear additionally for a moment. If you hit the observation limit, investigate the conditions and logs of the same run. Do not declare failure from a timeout alone or restart the installation. Immutable ConfigMaps can also be deleted and recreated, so check the preservation of the earlier UID and contents. Do not put secret values in a ConfigMap. Observation records are learning evidence, not remote attestation that stops the root user or an anti-cheating device. Preparation for earlier steps fills in only the earlier inputs and observations that are missing, and does not produce the current answer. It does not overwrite steps you have already done or a student's unfinished files. Grading of the current step also looks at the actual state, while past steps you have moved beyond use the saved records. After the final step, it keeps checking the recovery state and the other team's health. If you lose learner files, do not reconstruct the lost past state as a success record. You must leave a note that there is no past evidence for that run. Official documentation: ConfigMap · Deployment rollback.

Connect the baseline declaration and the response

Save /root/cnpa-release/baseline.json with capture 1. Check that the two actual Pods of the cnpa-rollback/checkout Deployment read configuration v1 and are ready, and that the Pod in the service response is owned by that Deployment.

Do not stop at a single service response; connect the Pod→ReplicaSet→Deployment ownership relationship and the individual readiness responses.

Distinguish a configuration change from existing environment variables

Write the v1 ConfigMap settings (namespace cnpa-rollback) in /root/cnpa-release/mutable-config.json. Its data is CONFIG_REVISION="v2", HEALTHY="false", and immutable=false. Apply it with act 2 and save env-unchanged.json with capture 2. The configuration has changed, but the UIDs of the two initial Pods and their actual v1 environment variables must be preserved.

Environment variables are received when the container starts. Look separately at the current contents of the configuration object and the values the existing processes read.

Distinguish a healthy service from a failed new deployment

In /root/cnpa-release/mutable-release.json, write a JSON patch that sets only labhub.io/release in spec.template.metadata.annotations to mutable-v2. Save rollout-incomplete.json with act 3 and capture 3. One new Pod must be v2 and not ready (503), the two existing ones must be v1 and healthy, and the service responds from the existing Pods.

A template annotation change creates a new rollout. With maxUnavailable=0, the existing healthy capacity is not reduced first because of a new Pod that is not ready.

The first rollback that looks like success

Roll back with act 4 to the previous revision observed at installation, and save apparent-rollback.json with capture 4. Check together that rollout status succeeds and the first two Pods are healthy again, and that the contents of settings are still the bad value.

What gets rolled back is the Pod template. Do not pin the revision number; use the value you observed first.

Confirm the recurrence on a replacement Pod

Replace one healthy Pod in your own lab with act 5 and save replacement-failed.json with capture 5. Check that the deleted UID is gone and that the Pod with the new UID receives the bad v2 configuration. One existing Pod must still respond normally.

Compare the UIDs of the first Pods with the UID of the newly born Pod. That one healthy Pod remains does not mean the whole recovery was maintained.

Bind immutable configuration to the deployment unit

In /root/cnpa-release/versioned-configs.json, write two ConfigMaps as a v1 List. In the cnpa-rollback namespace, settings-v1 is CONFIG_REVISION="v1" and HEALTHY="true", settings-v2 is CONFIG_REVISION="v2" and HEALTHY="false", and both are immutable=true. In versioned-release.json, write a patch that specifies the template annotation labhub.io/release=versioned-v1 and envFrom.configMapRef.name=settings-v1 for the container with name=app. Test applying it and the rejection of an immutable modification with act 6, and save versioned-ready.json with capture 6.

A ConfigMap needs apiVersion, kind, metadata, data, and immutable. The deployment patch is a strategic merge patch, not a whole Deployment document.

Deploy a bad version reference

In /root/cnpa-release/bad-release.json, write a patch that specifies the template annotation labhub.io/release=versioned-v2 and envFrom.configMapRef.name=settings-v2 for the container with name=app. Save versioned-incomplete.json with act 7 and capture 7. The two existing immutable v1 Pods must be healthy, and the one new v2 Pod must fail to become ready.

Immutable configuration is not fixed by editing its contents; you reference a new name. This step is an experiment that deliberately selects a bad version.

Recovery that holds through the next Pod

With act 8, recover to the actual revision from step 6 and then replace one healthy Pod. Save replacement-ready.json with capture 8. Check that the replaced new UID is also healthy with the immutable v1 configuration, and that the UID of settings-v1 and the UID and HTTP response of the other team's cnpa-rollback-other/sentinel are preserved.

Even after recovery, do not swap in a new configuration object with the same name. Look at the preserved UID and the replacement Pod's actual response together.