TT Lab
Get started
Learn Learning paths Courses

Helm Deployment and Rollback Scenarios

Hooks Live Outside the Release

Continue in TT Lab

Summary in one line

Hook resources are not managed by the release. So they do not disappear even if you roll back, and if you do not attach a deletion policy they pile up in the cluster.

Why this is needed

The most common way to put a DB migration into a deployment is a pre-upgrade hook Job. But this Job is not part of the release. Helm applies the hook, waits for its completion, and then applies the main manifest. The Job object created by a hook does not go into the release Secret.

There are two results.

How it works

metadata:
  annotations:
    "helm.sh/hook": pre-upgrade,pre-install
    "helm.sh/hook-weight": "-5"          # 작을수록 먼저
    "helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded

The three values of hook-delete-policy mean different things.

Value When it deletes
before-hook-creation Right before creating the same hook in the next deployment (a failed Job is left so you can investigate)
hook-succeeded Immediately on success
hook-failed Immediately on failure (the logs disappear — usually a bad choice)

The basic form in practice is before-hook-creation,hook-succeeded. Successful ones are cleaned up, and failed ones are left until the next deployment so you can look at the logs.

Common misconceptions

The misconception that a failed hook causes a rollback. If a pre-upgrade hook fails, the upgrade is aborted and the release remains as failed. With --atomic a rollback is applied, but a migration that has already run does not come back.

What you should not do with hooks

Hooks are convenient, so you end up putting all sorts of things in them. There are three things you should not put in.

Long-running data migration. A job that moves millions of rows holds up the deployment for tens of minutes. During that time the release is locked in pending-upgrade and other deployments are blocked. Run this kind of work as a separate Job apart from the deployment, and make the application able to read both schemas.

Notifying external systems. Notifying Slack with a post-install hook is common, but if the hook fails, the whole deployment is shown as failed. There is no reason to treat a deployment as a failure because a notification did not go out. Put such things on the CI side.

Tests. helm test is a separate command and uses helm.sh/hook: test. If you mix it into the deployment process, the deployment is blocked every time a test wobbles.

If a hook gets stuck, the deployment stops

If a hook Job does not finish, helm upgrade keeps waiting. The default of --timeout is 5 minutes, and after that it treats it as a failure. But even after the timeout, the Job keeps running in the cluster. Helm only stops waiting; it does not kill the Job.

A dangerous situation arises here. If you judge it a failure and deploy again, the second migration starts while the earlier one is still running. If two ALTERs are applied to the same table at the same time, the whole service can stop from lock waits.

You block this in three ways.

spec:
  activeDeadlineSeconds: 240      # helm --timeout 5m 보다 짧게
  backoffLimit: 0
  template:
    spec:
      restartPolicy: Never

The order of hooks and ordinary manifests

When hooks and ordinary resources are mixed in one chart, the order gets confusing. The actual order is as follows.

pre-install 훅  →  일반 매니페스트 적용  →  post-install 훅
   (weight 순)        (Helm 의 종류별 순서)      (weight 순)

A trap people often step into here. A pre-install hook runs when the Secrets and ConfigMaps do not yet exist. If a hook Job references the release's Secret, it fails with "no such secret." The values a hook uses must either be created by the hook itself (a Secret in the same hook group given a low weight) or made in advance outside the release.

hook-weight is a string but is sorted numerically. If you give "10" and "9", 9 comes first. The convention of using negative numbers came from here — use -5 as the default and give -10 to things that must run even earlier.

What really matters in practice

Separate migrations that cannot be undone from deployment. A change that removes a column is deployed in three steps — ① add the new column (old code still works) ② deploy the new code ③ drop the old column. That way, whichever step you roll back at, both the previous and the next versions can tolerate it. This is called the expand-contract pattern.