TT Lab
Get started
Learn Learning paths Courses

Helm Deployment and Rollback Scenarios

What --atomic and --wait Actually Do

Continue in TT Lab

Summary in one line

--wait means "wait until everything is up," and --atomic means "roll back if everything cannot come up." If you use neither, a failed deployment is reported as a success.

Why this is needed

By default, helm upgrade sends the manifest to the API server and returns success right away. It does not look at whether the Pods actually come up. Even if you wrote the image tag wrong and get ImagePullBackOff, helm prints STATUS: deployed. The CI log is green, and only users suffer the outage.

If you add --wait, helm polls the ready state of Deployments, StatefulSets, and DaemonSets. If they are not ready within --timeout (default 5 minutes), it treats that as a failure. Now CI turns red.

But if you leave it in the failed state, a problem remains. The new revision has already been created and the resources are half changed. Here --atomic automatically invokes helm rollback to return to the previous state. --atomic includes --wait — naturally so, since if it does not wait it cannot know about the failure.

How it works

helm upgrade demo ./chart \
  --atomic \
  --timeout 5m \
  --history-max 10

This one line should be the basic form of a deployment script. And set the timeout value to the startup time of the slowest Pod + some margin. If it is too short, a perfectly healthy deployment is rolled back, and if it is too long, you learn about the outage late.

Common misconceptions

The misconception that things are safe if you have --atomic. atomic rolls back only Kubernetes resources. If a migration Job run as a hook has already changed the DB, that stays as it is. So schema changes are deployed in separate steps in a form that both the previous and next versions can tolerate (expand → deploy → clean up).

The misconception that raising the timeout solves it. ImagePullBackOff does not get better by waiting. A 30-minute timeout only tells you about the outage 30 minutes later.

When a release gets stuck in pending

If a deployment dies midway, the release remains as pending-upgrade or pending-install. In this state, the next deployment is rejected like this.

Error: another operation (install/upgrade/rollback) is in progress

Helm 3 stores release state in a Secret in the namespace, and there is no separate lock, so the "in progress" marker just stays. The process has already died and only the marker is left.

The order for freeing it is as follows.

helm history demo                    # 1) 마지막 리비전의 상태를 본다
helm rollback demo <직전 성공 리비전>  # 2) 대개 이걸로 풀린다

# 롤백도 거절되면 마지막 수단 — 갇힌 리비전의 Secret 을 지운다
kubectl -n <ns> get secret -l owner=helm,name=demo
kubectl -n <ns> delete secret sh.helm.release.v1.demo.v<갇힌 번호>

The last method deletes the record of that revision, so it cannot be undone. Before deleting, save the content with helm get manifest demo --revision <n> (with the revision number in the placeholder).

What a rollback cannot undo

helm rollback reverts the release manifest. Everything else stays as it is.

What comes back What does not come back
The spec of Deployments, Services, and ConfigMaps DB migrations run by hooks
Image tags, environment variables, replicas Data inside a PVC
Resource requests and limits Events and notifications sent to other systems
State held by a deleted resource

Be especially careful about the last row. If you remove one resource from the chart and deploy, Helm deletes it. A rollback creates it again, but what was inside is new. Removing a StatefulSet briefly and then putting it back is the path that leads to data loss.

The basic form of a deployment pipeline

helm upgrade --install demo ./chart   -f values/prod.yaml   --set image.tag="$GIT_SHA"   --atomic --timeout 5m --history-max 10   --wait-for-jobs

# helm 이 초록불을 준 뒤에 바깥에서 확인한다
curl -fsS --retry 5 --retry-delay 3 https://demo.example.com/healthz

Leaving out --wait-for-jobs is common. --wait waits for Deployments and StatefulSets but does not wait for Jobs. This is where the accident of new Pods receiving traffic before the migration Job has finished happens.

What really matters in practice

What --wait looks at is the ready state of the workload, not "the service is healthy." Even if the Pods are Ready, the responses may be 500. That is why the end of a deployment pipeline must always have a smoke test that hits it from the outside. That helm gave a green light and that users can use it are different propositions.