TT Lab
Get started
Learn Learning paths Courses

CNPA — Cloud Native Platform Engineering Associate

The rollback button does not restore configuration

Continue in TT Lab

In one sentence

Rolling a Deployment back to an earlier revision is different from restoring the earlier execution environment. Even if the rolled-back Pod template points to a ConfigMap of the same name, if the contents behind that name have changed, the next Pod is born with a different configuration.

Why this was needed

A new deployment of a payment service does not become ready. The operator rolls back to the previous revision. rollout status succeeds and the service responds. You think it is over, but a little later one Pod is replaced and the same failure appears again. Even though it was already rolled back, why does the same fault recur?

So that this question does not remain a mere matter of memory, I reproduced it on a real personal cluster. I pinned the image and the app code and changed only the configuration. The first two Pods read the healthy configuration through environment variables. Even after I changed the ConfigMap's contents to a bad value, the existing processes' environment variables stayed as they were. But a new Pod read the changed value from the same ConfigMap name and failed its readiness check.

Right after the first rollback, the old Pods were still alive and things looked normal. When I replaced one of those Pods, the new Pod received the bad configuration again. The command did not lie. The range we expected was wider than the range the command rolls back.

How it works

Separate the object's name, its contents, and the time it was read

envFrom.configMapRef.name in a Pod template is a name that points to a configuration object. It is not a copy or fingerprint of the configuration contents. A container that uses a ConfigMap as environment variables receives the values when it starts. The environment variables of an already running process do not change automatically just because you edited the ConfigMap.

So two Pods that came from the same Pod template can run with different configurations. The Pod born earlier has the pre-change value, and the Pod born later has the post-change value. The inference "the manifest is the same, so the execution is the same" is missing the time at which the configuration was read. How configuration is updated depends on how it is delivered, so do not apply the explanation for environment variables directly to volume mounts.

What the revision history contains

A Deployment revision is tied to changes in the Pod template. Changes to the image, the name of the configuration reference, and template labels or annotations can create a new rollout. Changing the contents of an external ConfigMap is not itself a change to the Deployment's Pod template.

When you go back to a previous revision, the previous template is selected. If that template still references settings, and the contents of settings are a bad value, the source of the problem remains. A deployment record that logs only revision numbers can hardly explain which configuration contents were actually running.

Include the configuration in the deployment artifact

In the comparison experiment, I split the healthy configuration into settings-v1 and the bad configuration into settings-v2. I attached immutable: true to each object, and I changed the Deployment's reference name according to the version. Now when you go back to the previous Pod template, the name of the configuration it references goes back with it.

The key is not the trick of appending a number to the name. It is preserving the configuration contents needed for the earlier run and linking the chosen configuration and the Pod template as a single deployment unit. In fact, the API rejected a request to modify the healthy immutable configuration. Then, after recovering from the bad version to the earlier reference, even when I replaced a Pod, the new Pod read the healthy configuration.

What it looks like in the field

When a team says "we promote the same image from development to production," recording only the image digest may be half the explanation. The configuration that varies by environment must also be reviewed and tracked. Even with the same image, behavior differs if the connection targets, feature flags, and readiness conditions differ. This example surfaces that difference as a small health value.

Immutable objects can also be deleted and recreated with the same name. So you also need delete permissions, the retention period for earlier configurations, and the change history of the deployment source. The names in the example are version names to aid understanding, not content-based addresses or signatures. No guarantee of tamper-proofing is made, even against administrators. Do not put secret values in a ConfigMap.

Also, even if you roll back the application template, database schemas and external requests that have already been processed do not go back. Before deployment, you must distinguish changes that can be rolled back from changes that need separate recovery. A platform's recovery button should not be a button that hides that boundary, but an interface that shows the necessary evidence and the limits.

What to check in the next reading

We look at why the evidence that the service responds differs from the evidence that the new deployment is complete. While the existing Pods receive traffic, we read the state in which the new Pod is broken, and establish criteria that extend, after recovery, to the replacement Pod as well.

Official documentation