Helm Deployment and Rollback Scenarios
Confirm That Hooks Live Outside the Release
Goal
You confirm by hand the fact that hook resources live outside the release. You attach the three annotations, place the two deletion policies side by side, work out the execution order, and even see that what a hook left behind stays as it is after a rollback.
Why it matters
The most common way to put a DB migration into a deployment is a pre-upgrade hook. But objects created
by a hook do not go into the release Secret. There are two results. What the hook did
remains even after a rollback, and if you do not attach a deletion policy, objects keep piling up.
Practical rules follow from here. Clean up successful hooks and leave failed hooks until the next deployment so you can look at the logs; split schema changes that cannot be undone apart from the deployment and do them in several steps; and give Job hooks their own upper limit shorter than helm's timeout. The last is because even when helm stops waiting, the Job keeps running.
Environment
This Pod brings up a real kube-apiserver with kwok, but containers are not actually
run. So a Job hook never finishes and leaves the release stuck. The first six steps
handle hooks with ConfigMaps, which are judged ready immediately, and the Job is turned off with a single
value in step 7 while you look only at the render result. The working directory is /root/hs-hook and the outputs go under
/root/hs-hook/out.
Steps
- Create a chart and attach the three hook annotations to
hook-migrate.yaml. - Install it as
payinhook-laband look at both the cluster and the manifest. - Add
hook-scratch.yamland compare the two deletion policies. - Add
hook-early.yamland write the execution order in/root/hs-hook/out/order.txt. - Make a
testhook intests/smoke.yamland confirm that it does not run at deployment time. - Save a
--no-hooksrender to/root/hs-hook/out/nohooks.yaml. - Make a Job hook with safeguards in
hook-migrate-job.yaml, leaving the default off. - Upgrade and roll back, then summarize in four lines in
/root/hs-hook/out/residue.txt.
Notes
- Annotation values are strings. If you write
hook-weightwithout quotes, it is rendered as a number and not recognized as a hook. - In step 7, do not upgrade with
migrationJob.enabledturned on. The Job does not finish, so the release gets stuck inpending-upgrade. - A common mistake is putting
hook-failedin the deletion policy. The moment it fails, the object is deleted and the logs you would use to find the cause disappear.
Attach the three hook annotations
In /root/hs-hook, create a chart with helm create svc and delete the svc/templates/tests directory. Then create a {{ .Release.Name }}-migrate ConfigMap in svc/templates/hook-migrate.yaml and write helm.sh/hook as pre-install,pre-upgrade, helm.sh/hook-weight as -5, and helm.sh/hook-delete-policy as before-hook-creation.
A hook is not a separate resource kind but an ordinary resource with annotations attached. In practice you use a Job, but this cluster does not actually run containers, so a Job hook never finishes. That is why the first six steps handle hooks with ConfigMaps, which are judged ready immediately. Annotation values must be strings, so wrap the weight in quotes.
The hook is in the cluster but not in the manifest
Install the chart into the hook-lab namespace under the name pay, and look for the hook ConfigMap in two places, kubectl and helm get manifest.
It is helm install pay ./svc -n hook-lab --create-namespace. When the installation finishes, the hook ConfigMap shows up in kubectl -n hook-lab get cm, but it is not in helm get manifest pay -n hook-lab. This is because hook resources are not managed by the release. This one fact is the whole reason a rollback cannot undo a hook.
Place the two deletion policies side by side and see the difference
Create a {{ .Release.Name }}-scratch ConfigMap in svc/templates/hook-scratch.yaml. The hook timing is the same as before, the weight is 0, and the deletion policy is hook-succeeded. Then do one upgrade.
hook-succeeded deletes the object as soon as it succeeds, and before-hook-creation keeps it until right before the same hook is created in the next deployment. After the upgrade finishes, if you look at kubectl -n hook-lab get cm, only one of the two remains. Whether you can look at the logs of a failed hook is decided here.
Decide the execution order with weight
Add a {{ .Release.Name }}-early ConfigMap hook to svc/templates/hook-early.yaml with weight -10. Then work out the execution order of the hooks that run in the pre-upgrade stage and write only the names, one per line, in /root/hs-hook/out/order.txt.
hook-weight is a string but is sorted numerically. The smaller runs first, and when the values are equal, the tie is broken by resource kind and name. That is why in practice you set the default to -5 and give -10 to things that must run even earlier. Do not guess the order; pull the annotations out of the helm template result, sort them, and check.
A test hook does not run at deployment time
Create a {{ .Release.Name }}-smoke Pod in svc/templates/tests/smoke.yaml with helm.sh/hook set to test and the deletion policy set to hook-succeeded. Then do one more upgrade.
A test hook does not step into the deployment process and runs only when you call helm test. So even after the upgrade finishes, that Pod is not in the cluster and is not in helm get manifest either. It does appear in the render result. Checking the difference among these three for yourself is this step. If you mix tests into the deployment process, the deployment is blocked every time a test wobbles, so the places were separated by design.
See what --no-hooks leaves out
Save the result of rendering helm template with --no-hooks to /root/hs-hook/out/nohooks.yaml. Only the hooks must be gone and the main manifest must stay as it is.
--no-hooks can be attached to helm install, upgrade, and template alike. You use it during incident response when you need to skip hooks and push only the resources through, but since it amounts to skipping the migration, you must use it knowing the consequences. After saving, check that there is no resource with a hook attached and that the count of the rest is unchanged.
Put safeguards on the migration Job
Create a pre-upgrade Job hook in svc/templates/hook-migrate-job.yaml. Turn it on and off with migrationJob.enabled in values.yaml, and the default must be false. Give the Job activeDeadlineSeconds (under 300), backoffLimit: 0, and restartPolicy: Never, and set the deletion policy to before-hook-creation,hook-succeeded.
The default of helm's --timeout is 5 minutes, but even after the timeout, the Job keeps running in the cluster. helm only stops waiting. So set activeDeadlineSeconds shorter than that so the Job finishes by itself first. Do not put hook-failed in the deletion policy, since the logs disappear the moment it fails. On this cluster a Job hook never finishes, so the default must stay off. If you upgrade with it on, the release gets stuck in pending.
What the hook did remains even after a rollback
Do one more upgrade, then go back with helm rollback, and check that the hook ConfigMaps are still there. Write the result of your check in four lines in /root/hs-hook/out/residue.txt. They are HOOK_ROLLED_BACK, HOOK_IN_MANIFEST, MIGRATION_UNDONE, and SAFE_PATTERN.
A rollback reapplies the objects the release manages from the old manifest. What the hooks created is not in that manifest, so it is not touched. So the schema the migration changed also stays as it is. For the first three lines, write yes or no as yes or no, and on the last line write the name of the pattern for safely shipping a schema change that cannot be undone. It is the one the reading called expand and contract.