Helm Deployment and Rollback Scenarios
Deploy It, Break It, Roll It Back
Goal
You make a deployment something that can be undone. You go one full cycle: deploy, break it on purpose, undo it in two ways, and finally pull out a stuck release.
Why it matters
The success or failure of deployment automation is decided not when it succeeds but by what state it is left in when it fails.
By default helm upgrade returns success without looking at whether Pods come up.
That creates a situation where CI is green and only users suffer an outage.
In this lab you create that situation yourself. In step 4, when you deploy with an image that does not exist, you will see with your own eyes that
helm says STATUS: deployed. Then you compare what changes when you attach --atomic.
Environment
This Pod brings up a real kube-apiserver with kwok. Pods do not actually run, but release Secrets, revisions, and resource changes are all real. So grading is also done by re-reading the cluster state, not files.
kubectl get nodes 노드 2대가 Ready
helm version v3.16
Steps
helm create demo→helm install- Pull the manifest out of the release Secret to
/root/helm/02-manifest.yaml - Revision 2 with
--set replicaCount=4 - Upgrade with an image that does not exist — confirm that it says success
- Repeat the same mistake with
--atomic→/root/helm/05-atomic.txt helm rollback demo 2- Diff two runs of
helm template→/root/helm/07-diff.txt - Pull out a release stuck in pending-upgrade
Notes
- In this lab there are steps where the command failing is also the correct answer (step 5). The purpose is to save the failure output.
- Look at
helm history demooften. Everything that happened is there. - Stuck releases arise in practice when the helm process dies midway or CI is canceled.
Create a chart and deploy it
helm create demo → helm install
Create the skeleton with helm create demo and deploy with helm install demo ./demo. It is fine if it shows as deployed in helm list. This cluster is kwok, so Pods do not actually run, but the releases, revisions, and resources are all real.
Look directly at the Secret where the release is stored
Pull the manifest out of the release Secret to /root/helm/02-manifest.yaml
Check the name with kubectl get secret -l owner=helm. If you decode in the order -o jsonpath='{.data.release}' → base64 -d → base64 -d → gzip -d, the whole release JSON comes out (not the manifest). The rendered YAML is in the .manifest field of that JSON as a string, so pull it out with jq -r .manifest and save it to /root/helm/02-manifest.yaml.
Raise replicas to make revision 2
Revision 2 with --set replicaCount=4
helm upgrade demo ./demo --set replicaCount=4. Then check with helm history demo whether there are now two revisions and whether number 1 became superseded.
Break it on purpose
Upgrade with an image that does not exist — confirm that it says success
Try upgrading with an image that does not exist: --set image.repository=nope/nothing --set image.tag=v0. Without --wait, helm says it succeeded — that is what you are to see in this step. You will undo it in later steps, so grading looks at the revision record, not the current state.
Catch the failure with --atomic
Repeat the same mistake with --atomic → /root/helm/05-atomic.txt
This time, create a Pod that cannot be scheduled: helm upgrade demo ./demo --set nodeSelector.disktype=nope --atomic --timeout 30s. Because it is a node label that does not exist, the Pod stays Pending, and --atomic catches the timeout and rolls back automatically. The command failing is the correct answer — save the output to /root/helm/05-atomic.txt (2>&1 | tee).
Why nodeSelector rather than a broken image: the cluster in this lab is kwok, which does not actually run Pods and marks them Ready. So even if the image is wrong, --wait passes. It truly stops only with a condition where scheduling itself is impossible.
Undo it by hand
helm rollback demo 2
Go back to the content of revision 2 with helm rollback demo 2. A new revision with 'Rollback to 2' written on it appears in helm history, and kubectl get deploy demo -o jsonpath='{.spec.replicas}' must be 4.
Preview what will change
Diff two runs of helm template → /root/helm/07-diff.txt
The habit of looking at the difference before applying prevents accidents. The helm-diff plugin is not available, so compare the results of two helm template runs with diff. Save the result to /root/helm/07-diff.txt. There must be a difference.
Pull out a release stuck in pending
Pull out a release stuck in pending-upgrade
First create the stuck situation for real. Have a Pod that cannot be scheduled wait under --wait, and kill that helm process midway.
helm upgrade demo ./demo --set nodeSelector.disktype=nope --wait --timeout 300s &
sleep 12
kill -9 $!
In practice, when the helm process dies midway or CI is canceled, you end up in exactly this state. You check with helm list -a — the default of helm list hides pending. If you run helm upgrade again in this state, you are blocked with another operation (install/upgrade/rollback) is in progress.
(Attaching just kubectl label ... status=pending-upgrade to a Secret does not get you stuck. helm reads state not from labels but from the release JSON inside the Secret, so if you change only the label, helm list still answers deployed and the next deployment goes out as usual.)
The procedure for pulling it out: helm does not clean up a stuck revision Secret on its own. Find it with kubectl get secret -l owner=helm,name=demo,status=pending-upgrade, delete it with kubectl delete, and then helm rollback to the last healthy revision. When you are done, helm list must show deployed and there must be no Secret carrying the pending label.