TT Lab
Get started
Learn Learning paths Courses

Authoring and Shipping Helm Charts

A Release — The Deployment's Memory Kept Inside the Cluster

Continue in TT Lab

Summary in one line

A release is not a record of commands but state stored inside the cluster, and that state accumulates as one Secret per revision.

Why this is needed

When you deploy with kubectl apply, you can tell "what is running now" but not "what was running yesterday." To roll back when something goes wrong, you have to fetch the previous manifest from somewhere, and in most cases that somewhere is a person's memory or someone's laptop.

Helm solves this problem with "leave a full snapshot of each deployment in the cluster every time you deploy." What remains is not just the rendered manifest. The chart metadata used at the time, the values the user passed, the status, and the release notes all go in whole. So a rollback becomes not "finding an old file" but "choosing a stored revision."

How it works

Helm 3 has no server component resident in the cluster. The CLI talks directly to the API server with a kubeconfig, so Kubernetes RBAC applies as it is. Then where does it keep state — it keeps it in a Secret in the namespace where the release is installed.

Item Value
Secret name sh.helm.release.v1.<릴리스이름>.v<리비전번호> (release name and revision number)
Secret type helm.sh/release.v1
Labels owner=helm, name=<릴리스이름>, status=<상태> (release name and status)
Content Chart metadata, rendered manifest, values, status, and notes, gzip-compressed and then base64-encoded

One Secret is added per revision. So just looking at kubectl get secret -l owner=helm shows how many layers of deployment history that namespace has. By default it keeps up to the latest 10, adjustable with --history-max.

An upgrade is not a simple overwrite. Helm 3 compares three things — the manifest of the previous revision, the actual current state in the cluster, and the newly rendered manifest. This is a three-way strategic merge patch. Thanks to it, it notices fields that someone touched directly with kubectl, leaves fields that the chart does not manage alone, and applies only what changed.

And there are two properties you must remember.

First, a failed upgrade also remains as a revision. If rendering succeeded but the API server rejected it, that revision is recorded in the failed state, and the real objects stay as they were before. A failure remaining in the history is not an accident but a feature — what was tried and why it did not work stays in the cluster.

Second, a rollback does not undo; it creates a new revision. If you roll back to revision 2, the number does not return to 2; a new revision 4, computed from revision 2's chart and values, is created. Revision numbers only ever move forward. Thanks to this property, even "rolling back and then canceling the rollback" remains intact in the history.

There are two ways to look up values as well. helm get values shows only the values the user actually passed, and with -a (--all) it shows the final values merged with the chart defaults. If you also give --revision N, you can see the values of a specific revision. What separates "is this setting a default or a value someone put in?" during incident response is the difference between these two commands.

What it looks like in the field

First, the stuck pending state. If the process dies during an upgrade, the release remains as pending-upgrade and the next command is refused. What you need then is not to delete the Secret by hand but to roll back to the last successful revision and clean up the state.

Second, the two faces of --atomic. --atomic includes --wait and automatically rolls back on failure, which is good for pipelines. But once it automatically rolls back, the chance to observe the failed state is gone. In situations where you need to see the cause, deliberately do not attach it.

Third, value drift. If someone quickly raised the replica count with --set to get through an outage, that value exists only in that revision. The next deployment quietly reverts it. So values added temporarily must always be reflected in the value file in the repository.

The order for freeing a stuck release

Helm is convenient when deployments go well, and once something goes off, you have to fix the state by hand. There are three you often see.

Stuck in pending-upgrade. If CI dies or times out during an upgrade, the release stays in that state, and the next helm upgrade is rejected with "another operation is in progress." After confirming that there is no job actually running, roll back.

helm history myapp -n prod
helm rollback myapp <마지막 deployed 리비전> -n prod

--atomic and --wait do different things. --wait waits until the resources are ready, and --atomic automatically rolls back if it fails while waiting. In a pipeline it is better to use --atomic --timeout 10m together. Otherwise a half-applied state is left behind.

Hooks create deadlocks. Suppose a pre-upgrade hook runs a migration job, that job uses the new version's image, and that image gets deployed only after the upgrade finishes — then it waits forever. For hooks, use an image that already exists, and use hook-delete-policy to keep the failed job so you can look at its logs.

annotations:
  "helm.sh/hook": pre-upgrade
  "helm.sh/hook-weight": "-5"
  "helm.sh/hook-delete-policy": before-hook-creation

Check where values came from. When several -f files and --set are mixed, the final value gets confusing. The one that comes later wins, and --set is stronger than files.

helm get values myapp -n prod --all      # 실제로 쓰인 값(기본값 포함)
helm template . -f values-prod.yaml | kubectl diff -f -

CRDs are not updated on upgrade. What is in the crds/ directory is applied only at install. If you upgraded the chart but a new field does not take effect, this is the spot. CRDs must be applied separately.

What you will do in the next lab

You install the release lab-app in the namespace helm-lab to make revision 1, and raise the replica count to make revision 2. Then you deliberately put in a value the API server will reject to make failed revision 3, and roll back to revision 2 to get revision 4. You extract helm get values with and without -a and compare them, directly check the release Secrets accumulated in the namespace, and then produce a lifecycle report.