TT Lab
Get started
Learn Learning paths Courses

Helm Deployment and Rollback Scenarios

A Release Is One Secret

Continue in TT Lab

Summary in one line

A Helm release is not magic but a single Secret in a namespace. Once you know that, it explains everything about what a rollback undoes and why some things cannot be undone.

Why this is needed

It is common to have run helm rollback only to find that the database migration did not come back and the service broke. Conversely, the report "I rolled back but nothing changed" is also common. Both come from the same misunderstanding — thinking that a rollback rewinds time.

What Helm does is much simpler. When you run helm install, it compresses the entire rendered manifest with gzip and stores it in a Secret. The name is sh.helm.release.v1.<릴리스>.v<리비전> (release and revision). When you run helm upgrade, it creates one more revision Secret. helm rollback 3 takes out the Secret of revision 3 and applies its manifest again.

So the rules follow. What was written in the manifest comes back. What happened outside the manifest — data changed by a migration, files created by a Job, requests sent to an external API — does not come back.

How it works

You can check this directly.

kubectl get secret -l owner=helm
kubectl get secret sh.helm.release.v1.demo.v1 -o jsonpath='{.data.release}' \
  | base64 -d | base64 -d | gzip -d | head -40

Decoding base64 twice looks strange, but it is right. Kubernetes encodes the Secret value once, and Helm had already encoded it once more inside.

helm history is a list of these Secrets. Each revision has a status.

Status Meaning
deployed The revision that is alive now. There is always only one
superseded One that was deployed earlier and gave way to the next revision
failed A revision that failed to apply
pending-upgrade An upgrade that started and did not finish. If you get stuck here, the next deployment is blocked

Common misconceptions

A rollback does not turn the revision number back. If you run helm rollback demo 1, you do not go back to revision 1; instead, a new revision 3 is created with the content of revision 1. If you look at the history, line 3 says Rollback to 1. This is a design to avoid erasing the audit record — what happened has to remain in the history.

There is a limit on how many are kept. The default is 10 (--history-max). Old revisions are deleted, so "rolling back to six months ago" is usually impossible. The window you can go back to is narrower than you think.

Looking directly inside a release Secret

Helm 3 creates one Secret per release. If you know this structure, you can recover by hand during an incident.

kubectl -n labhub-prod get secret -l owner=helm,name=labhub
NAME                             TYPE                 DATA   AGE
sh.helm.release.v1.labhub.v247   helm.sh/release.v1   1      2d
sh.helm.release.v1.labhub.v248   helm.sh/release.v1   1      1d
sh.helm.release.v1.labhub.v249   helm.sh/release.v1   1      3h

# 안을 풀어 본다 — base64 → gzip → JSON
kubectl get secret sh.helm.release.v1.labhub.v249 -o jsonpath='{.data.release}'   | base64 -d | base64 -d | gzip -d | jq '.info, .chart.metadata.version'

The base64 being applied twice is the trap. The Kubernetes Secret itself is one layer, and Helm adds another layer when it stores.

Two conclusions come out of this.

Telling the three states apart

helm list          → deployed 만 보인다
helm list --all    → failed, pending-upgrade, superseded 까지
helm history <릴리스> → 리비전별 상태와 설명

superseded is normal — it is something that stepped aside because a new revision came out. The problem is pending-*. If it stays in this state, the next deployment is refused, and the process has already died with only the marker left behind.

You must not leave failed as it is either. The next deployment will work, but the reference point for the --atomic rollback gets blurry. Fix the cause and make one successful deployment to tidy up.

Release names and resource names

A chart's resource names are usually built as {{ .Release.Name }}-{{ .Chart.Name }}. So if you change the release name, all the resources are created anew — the old ones remain and new ones appear, so the two coexist.

For the same reason, choose the release name carefully at first, and if you have to change it, plan a procedure to delete the old release and install a new one. For a workload that uses a PVC, a step to move the data goes in between.

What really matters in practice

That a release is a Secret also means there is a size limit. etcd's object limit is 1MiB. If the chart grows (especially if it holds a lot of CRDs), the release Secret runs into that limit and upgrades fail. The error message that comes out then gives no hint of the cause, so knowing this structure is exactly what determines the speed of the fix.