TT Lab
Get started
Learn Learning paths Courses

Helm Deployment and Rollback Scenarios

Deploy It, Break It, Roll It Back

Continue in TT Lab

Goal

You make a deployment something that can be undone. You go one full cycle: deploy, break it on purpose, undo it in two ways, and finally pull out a stuck release.

Why it matters

The success or failure of deployment automation is decided not when it succeeds but by what state it is left in when it fails. By default helm upgrade returns success without looking at whether Pods come up. That creates a situation where CI is green and only users suffer an outage.

In this lab you create that situation yourself. In step 4, when you deploy with an image that does not exist, you will see with your own eyes that helm says STATUS: deployed. Then you compare what changes when you attach --atomic.

Environment

This Pod brings up a real kube-apiserver with kwok. Pods do not actually run, but release Secrets, revisions, and resource changes are all real. So grading is also done by re-reading the cluster state, not files.

kubectl get nodes            노드 2대가 Ready
helm version                 v3.16

Steps

  1. helm create demo → helm install
  2. Pull the manifest out of the release Secret to /root/helm/02-manifest.yaml
  3. Revision 2 with --set replicaCount=4
  4. Upgrade with an image that does not exist — confirm that it says success
  5. Repeat the same mistake with --atomic → /root/helm/05-atomic.txt
  6. helm rollback demo 2
  7. Diff two runs of helm template → /root/helm/07-diff.txt
  8. Pull out a release stuck in pending-upgrade

Notes

Create a chart and deploy it

helm create demo → helm install

Create the skeleton with helm create demo and deploy with helm install demo ./demo. It is fine if it shows as deployed in helm list. This cluster is kwok, so Pods do not actually run, but the releases, revisions, and resources are all real.

Look directly at the Secret where the release is stored

Pull the manifest out of the release Secret to /root/helm/02-manifest.yaml

Check the name with kubectl get secret -l owner=helm. If you decode in the order -o jsonpath='{.data.release}' → base64 -d → base64 -d → gzip -d, the whole release JSON comes out (not the manifest). The rendered YAML is in the .manifest field of that JSON as a string, so pull it out with jq -r .manifest and save it to /root/helm/02-manifest.yaml.

Raise replicas to make revision 2

Revision 2 with --set replicaCount=4

helm upgrade demo ./demo --set replicaCount=4. Then check with helm history demo whether there are now two revisions and whether number 1 became superseded.

Break it on purpose

Upgrade with an image that does not exist — confirm that it says success

Try upgrading with an image that does not exist: --set image.repository=nope/nothing --set image.tag=v0. Without --wait, helm says it succeeded — that is what you are to see in this step. You will undo it in later steps, so grading looks at the revision record, not the current state.

Catch the failure with --atomic

Repeat the same mistake with --atomic → /root/helm/05-atomic.txt

This time, create a Pod that cannot be scheduled: helm upgrade demo ./demo --set nodeSelector.disktype=nope --atomic --timeout 30s. Because it is a node label that does not exist, the Pod stays Pending, and --atomic catches the timeout and rolls back automatically. The command failing is the correct answer — save the output to /root/helm/05-atomic.txt (2>&1 | tee).

Why nodeSelector rather than a broken image: the cluster in this lab is kwok, which does not actually run Pods and marks them Ready. So even if the image is wrong, --wait passes. It truly stops only with a condition where scheduling itself is impossible.

Undo it by hand

helm rollback demo 2

Go back to the content of revision 2 with helm rollback demo 2. A new revision with 'Rollback to 2' written on it appears in helm history, and kubectl get deploy demo -o jsonpath='{.spec.replicas}' must be 4.

Preview what will change

Diff two runs of helm template → /root/helm/07-diff.txt

The habit of looking at the difference before applying prevents accidents. The helm-diff plugin is not available, so compare the results of two helm template runs with diff. Save the result to /root/helm/07-diff.txt. There must be a difference.

Pull out a release stuck in pending

Pull out a release stuck in pending-upgrade

First create the stuck situation for real. Have a Pod that cannot be scheduled wait under --wait, and kill that helm process midway.

helm upgrade demo ./demo --set nodeSelector.disktype=nope --wait --timeout 300s &
sleep 12
kill -9 $!

In practice, when the helm process dies midway or CI is canceled, you end up in exactly this state. You check with helm list -a — the default of helm list hides pending. If you run helm upgrade again in this state, you are blocked with another operation (install/upgrade/rollback) is in progress.

(Attaching just kubectl label ... status=pending-upgrade to a Secret does not get you stuck. helm reads state not from labels but from the release JSON inside the Secret, so if you change only the label, helm list still answers deployed and the next deployment goes out as usual.)

The procedure for pulling it out: helm does not clean up a stuck revision Secret on its own. Find it with kubectl get secret -l owner=helm,name=demo,status=pending-upgrade, delete it with kubectl delete, and then helm rollback to the last healthy revision. When you are done, helm list must show deployed and there must be no Secret carrying the pending label.