Helm Deployment and Rollback Scenarios
Hooks Live Outside the Release
Summary in one line
Hook resources are not managed by the release. So they do not disappear even if you roll back, and if you do not attach a deletion policy they pile up in the cluster.
Why this is needed
The most common way to put a DB migration into a deployment is a pre-upgrade hook Job. But this Job is not part of the release. Helm applies the hook, waits for its completion, and then applies the main manifest. The Job object created by a hook does not go into the release Secret.
There are two results.
- What the hook did remains even after a rollback. The schema the migration changed stays as it is.
- Job objects keep piling up. If you do not attach
helm.sh/hook-delete-policy, one more is added on every deployment.
How it works
metadata:
annotations:
"helm.sh/hook": pre-upgrade,pre-install
"helm.sh/hook-weight": "-5" # 작을수록 먼저
"helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded
The three values of hook-delete-policy mean different things.
| Value | When it deletes |
|---|---|
before-hook-creation |
Right before creating the same hook in the next deployment (a failed Job is left so you can investigate) |
hook-succeeded |
Immediately on success |
hook-failed |
Immediately on failure (the logs disappear — usually a bad choice) |
The basic form in practice is before-hook-creation,hook-succeeded. Successful ones are cleaned up, and failed ones are left until the next deployment so you can look at the logs.
Common misconceptions
The misconception that a failed hook causes a rollback. If a pre-upgrade hook fails, the upgrade is aborted and the release remains as failed. With --atomic a rollback is applied, but a migration that has already run does not come back.
What you should not do with hooks
Hooks are convenient, so you end up putting all sorts of things in them. There are three things you should not put in.
Long-running data migration. A job that moves millions of rows holds up the deployment for tens of minutes.
During that time the release is locked in pending-upgrade and other deployments are blocked. Run this kind of work as a
separate Job apart from the deployment, and make the application able to read both schemas.
Notifying external systems. Notifying Slack with a post-install hook is common, but
if the hook fails, the whole deployment is shown as failed. There is no reason to
treat a deployment as a failure because a notification did not go out. Put such things on the CI side.
Tests. helm test is a separate command and uses helm.sh/hook: test. If you mix it into the
deployment process, the deployment is blocked every time a test wobbles.
If a hook gets stuck, the deployment stops
If a hook Job does not finish, helm upgrade keeps waiting. The default of --timeout is
5 minutes, and after that it treats it as a failure. But even after the timeout, the Job
keeps running in the cluster. Helm only stops waiting; it does not kill the Job.
A dangerous situation arises here. If you judge it a failure and deploy again, the second migration starts while the earlier one is still running. If two ALTERs are applied to the same table at the same time, the whole service can stop from lock waits.
You block this in three ways.
- Set
activeDeadlineSecondson the Job. If you set it shorter than Helm's timeout, the Job finishes by itself first. backoffLimit: 0. For migrations, retrying is often not safe.- Use the migration tool's lock. Flyway, Liquibase, and Alembic all have a lock table, so
a second run waits. For a home-made script, use
pg_advisory_lock.
spec:
activeDeadlineSeconds: 240 # helm --timeout 5m 보다 짧게
backoffLimit: 0
template:
spec:
restartPolicy: Never
The order of hooks and ordinary manifests
When hooks and ordinary resources are mixed in one chart, the order gets confusing. The actual order is as follows.
pre-install 훅 → 일반 매니페스트 적용 → post-install 훅
(weight 순) (Helm 의 종류별 순서) (weight 순)
A trap people often step into here. A pre-install hook runs when the Secrets and ConfigMaps do not yet exist.
If a hook Job references the release's Secret, it fails with "no such secret." The values a hook uses must either be created by the hook itself (a Secret
in the same hook group given a low weight) or made in advance outside the release.
hook-weight is a string but is sorted numerically. If you give "10" and "9",
9 comes first. The convention of using negative numbers came from here — use -5 as the default
and give -10 to things that must run even earlier.
What really matters in practice
Separate migrations that cannot be undone from deployment. A change that removes a column is deployed in three steps — ① add the new column (old code still works) ② deploy the new code ③ drop the old column. That way, whichever step you roll back at, both the previous and the next versions can tolerate it. This is called the expand-contract pattern.