TT Lab
Get started
Learn Learning paths Courses

CKAD — Kubernetes Application Developer

Measure Whether the Deployment Is Really Zero-Downtime

Continue in TT Lab

This lab runs on a cluster where Pods really start and die

A real k3s is running inside a VM. A rolling update really replaces Pods one at a time, a failed rollout really stops, and kubectl drain really gets blocked by a PodDisruptionBudget.

In the fake cluster that the other CKAD rollout labs run on, Pods do not run, so there is no way to measure whether it was "zero downtime."

The first startup takes about 2 minutes.

The estimate is 70 minutes. The session defaults to 60 minutes, so extend it with +time before it expires (up to 180 minutes). When the session ends, the VM and files are reclaimed, so keep any results you need separately.

Goal

You keep sending requests before, during, and after the Deployment replacement and actually count HTTP failures, and you see how a failed rollout stops and what a rollback actually reverts.

Why it matters

"Zero-downtime deployment" is not about choosing a strategy name; it is a result that holds only when several settings mesh together. If even one is missing, it quietly breaks.

And whether it broke or not, you only know by measuring. If you check after the deployment is over, it always looks normal.

Steps

Create everything in the crol namespace.

  1. Create the web Deployment (3 replicas, nginx, with a readinessProbe) and explicitly set maxSurge and maxUnavailable. Record it in /root/crol/rolling.txt.
  2. With web set to replicas=3, maxSurge=1, maxUnavailable=0 and a readinessProbe, observe two real rollouts with python3 /opt/fixtures/ckad_rollout_observer.py capture. The first Pods must have no preStop. In /root/crol/nodowntime.txt, write total, failed, total_before_prestop, and failed_before_prestop as key=value, and also write scope=single-vm-rollout-sample and hook_not_retroactive=yes. Record the failure counts as observed.
  3. Update to a nonexistent image, confirm that the rollout stops, and record in /root/crol/stuck.txt that the old Pods stay alive in the meantime.
  4. Roll back with rollout undo and record the result in /root/crol/undo.txt. Also set revisionHistoryLimit explicitly.
  5. Use rollout pause and resume to create a state where only some are replaced, and record it in /root/crol/pause.txt.
  6. Create the web-pdb PodDisruptionBudget, and record in /root/crol/pdb.txt how many can be disrupted right now and what the PDB cannot prevent.
  7. Create the legacy Deployment with the Recreate strategy, and record in /root/crol/recreate.txt why such a strategy is needed.
  8. In /root/crol/report.md, write three lines, downtime_requests_failed=, recreate_causes_downtime=, and pdb_min_available=, along with an explanation.

Notes

Two numbers set the replacement speed

Create the web Deployment (3 replicas, nginx, with a readinessProbe) and explicitly set maxSurge and maxUnavailable. Record it in /root/crol/rolling.txt.

maxSurge is how many more than the target may be started, and maxUnavailable is how many may be missing.

Measure whether it is zero downtime with numbers

Prepare web with replicas=3, maxSurge=1, maxUnavailable=0, and a readinessProbe. The first Pods must have no preStop. After observing two real rollouts with python3 /opt/fixtures/ckad_rollout_observer.py capture, write total, failed, total_before_prestop, and failed_before_prestop as key=value in /root/crol/nodowntime.txt. Also write scope=single-vm-rollout-sample and hook_not_retroactive=yes, and explain the result. Do not assume the failure count is 0; record it as observed.

The trials in rollout-observation.json are in the order without_hook, with_hook. Look together at the HTTP status 200 of each request, the nginx body, and whether there is an error. The hook installation rollout is finished separately from the measurement, so check whether the before Pods of the second measurement have the hook.

A failed rollout stops

Update to a nonexistent image, confirm that the rollout stops, and record in /root/crol/stuck.txt that the old Pods stay alive in the meantime.

If you update to a nonexistent image, the new Pod does not start and the rollout stops after progressDeadlineSeconds. Look at the old Pods in the meantime.

What does a rollback revert

Roll back with rollout undo and record the result in /root/crol/undo.txt. Also set revisionHistoryLimit explicitly.

rollout undo returns to the immediately previous version. The number of remaining versions is decided by revisionHistoryLimit.

Change only part and look

Use rollout pause and resume to create a state where only some are replaced, and record it in /root/crol/pause.txt.

If you change the image after rollout pause, nothing happens. Use it when you want to gather several changes and roll them out at once.

Keep them from all going down at once

Create the web-pdb PodDisruptionBudget, and record in /root/crol/pdb.txt how many can be disrupted right now and what the PDB cannot prevent.

A PDB blocks only voluntary disruptions. disruptionsAllowed tells you how many are allowed right now.

A strategy that interrupts on purpose

Create the legacy Deployment with the Recreate strategy, and record in /root/crol/recreate.txt why such a strategy is needed.

Recreate brings down all the old ones and then starts the new ones. Use it when two versions must never coexist.

What you learned

In /root/crol/report.md, write three lines, downtime_requests_failed=, recreate_causes_downtime=, and pdb_min_available=, along with an explanation.

Write three lines, downtime_requests_failed=, recreate_causes_downtime=, and pdb_min_available=, along with an explanation.