CKAD — Kubernetes Application Developer
Measure Whether the Deployment Is Really Zero-Downtime
This lab runs on a cluster where Pods really start and die
A real k3s is running inside a VM. A rolling update really replaces Pods one at a time,
a failed rollout really stops, and kubectl drain really gets blocked by a
PodDisruptionBudget.
In the fake cluster that the other CKAD rollout labs run on, Pods do not run, so there is no way to measure whether it was "zero downtime."
The first startup takes about 2 minutes.
The estimate is 70 minutes. The session defaults to 60 minutes, so extend it with +time before it expires (up to 180 minutes). When the session ends, the VM and files are reclaimed, so keep any results you need separately.
Goal
You keep sending requests before, during, and after the Deployment replacement and actually count HTTP failures, and you see how a failed rollout stops and what a rollback actually reverts.
Why it matters
"Zero-downtime deployment" is not about choosing a strategy name; it is a result that holds only when several settings mesh together. If even one is missing, it quietly breaks.
- Without a
readinessProbe, a new Pod receives traffic as soon as it starts, even though it is not ready yet. - If
maxUnavailableis large, many go down at once and capacity falls short. - If the termination grace period (
terminationGracePeriodSeconds) is short, requests in progress are cut off.
And whether it broke or not, you only know by measuring. If you check after the deployment is over, it always looks normal.
Steps
Create everything in the crol namespace.
- Create the
webDeployment (3 replicas, nginx, with a readinessProbe) and explicitly setmaxSurgeandmaxUnavailable. Record it in/root/crol/rolling.txt. - With web set to replicas=3, maxSurge=1, maxUnavailable=0 and a readinessProbe, observe two real rollouts with
python3 /opt/fixtures/ckad_rollout_observer.py capture. The first Pods must have no preStop. In/root/crol/nodowntime.txt, writetotal,failed,total_before_prestop, andfailed_before_prestopas key=value, and also writescope=single-vm-rollout-sampleandhook_not_retroactive=yes. Record the failure counts as observed. - Update to a nonexistent image, confirm that the rollout stops, and record in
/root/crol/stuck.txtthat the old Pods stay alive in the meantime. - Roll back with
rollout undoand record the result in/root/crol/undo.txt. Also setrevisionHistoryLimitexplicitly. - Use
rollout pauseandresumeto create a state where only some are replaced, and record it in/root/crol/pause.txt. - Create the
web-pdbPodDisruptionBudget, and record in/root/crol/pdb.txthow many can be disrupted right now and what the PDB cannot prevent. - Create the
legacyDeployment with theRecreatestrategy, and record in/root/crol/recreate.txtwhy such a strategy is needed. - In
/root/crol/report.md, write three lines,downtime_requests_failed=,recreate_causes_downtime=, andpdb_min_available=, along with an explanation.
Notes
- The helper in step 2 changes a template annotation on the same image to trigger real rollouts. The order is: a rollout without the hook → hook installation complete → a separate rollout of the Pods that have the hook. Do not interpret it as retroactively applying the new template's hook to the old Pods.
rollout-observation.jsonpreserves the real requests and the Pod UIDs and hook markers before and after replacement. HTTP 500, redirects, wrong bodies, and transport errors are counted as failures. Do not edit the observation file to make the report's values fit. A completed capture is not redeployed and the existing observation is preserved.- Write the results as observed even if the failures are 0 both times or there is no difference. This sample alone cannot prove the hook's effect or zero downtime for the whole service including external load balancers and long-lived connections. An interrupted observation is not overwritten automatically and is redone in a new session.
- Check the rollout status with
kubectl -n crol rollout status deploy/web, and the reason it stopped is in.status.conditions. The example answer uses a 60-second deadline for normal startup and a 15-second deadline only for the intentional image failure experiment, then returns to 60 seconds after recovery. This is a choice for a short teaching experiment, not a recommended value for production services. - The Correct view and step preparation do not overwrite an existing answer. If the existing answer is wrong, look at the failure reason and fix it yourself. If you run the two features at the same time, it waits for the command already running and then tries again.
rollout undocan pick a specific version with--to-revision, and you see the remaining versions withrollout history.- The PDB's current slack is
.status.disruptionsAllowed. If there are not enough Pods, it becomes 0 and the drain is blocked. - Common mistake 1: expecting zero downtime without a
readinessProbe. A new Pod enters the endpoints as soon as it starts and receives traffic before it is ready. - Common mistake 2: thinking a PDB also protects against node failure. It blocks only voluntary disruptions (drain, eviction). It cannot stop a node from suddenly dying.
Two numbers set the replacement speed
Create the web Deployment (3 replicas, nginx, with a readinessProbe) and explicitly set maxSurge and maxUnavailable. Record it in /root/crol/rolling.txt.
maxSurge is how many more than the target may be started, and maxUnavailable is how many may be missing.
Measure whether it is zero downtime with numbers
Prepare web with replicas=3, maxSurge=1, maxUnavailable=0, and a readinessProbe. The first Pods must have no preStop. After observing two real rollouts with python3 /opt/fixtures/ckad_rollout_observer.py capture, write total, failed, total_before_prestop, and failed_before_prestop as key=value in /root/crol/nodowntime.txt. Also write scope=single-vm-rollout-sample and hook_not_retroactive=yes, and explain the result. Do not assume the failure count is 0; record it as observed.
The trials in rollout-observation.json are in the order without_hook, with_hook. Look together at the HTTP status 200 of each request, the nginx body, and whether there is an error. The hook installation rollout is finished separately from the measurement, so check whether the before Pods of the second measurement have the hook.
A failed rollout stops
Update to a nonexistent image, confirm that the rollout stops, and record in /root/crol/stuck.txt that the old Pods stay alive in the meantime.
If you update to a nonexistent image, the new Pod does not start and the rollout stops after progressDeadlineSeconds. Look at the old Pods in the meantime.
What does a rollback revert
Roll back with rollout undo and record the result in /root/crol/undo.txt. Also set revisionHistoryLimit explicitly.
rollout undo returns to the immediately previous version. The number of remaining versions is decided by revisionHistoryLimit.
Change only part and look
Use rollout pause and resume to create a state where only some are replaced, and record it in /root/crol/pause.txt.
If you change the image after rollout pause, nothing happens. Use it when you want to gather several changes and roll them out at once.
Keep them from all going down at once
Create the web-pdb PodDisruptionBudget, and record in /root/crol/pdb.txt how many can be disrupted right now and what the PDB cannot prevent.
A PDB blocks only voluntary disruptions. disruptionsAllowed tells you how many are allowed right now.
A strategy that interrupts on purpose
Create the legacy Deployment with the Recreate strategy, and record in /root/crol/recreate.txt why such a strategy is needed.
Recreate brings down all the old ones and then starts the new ones. Use it when two versions must never coexist.
What you learned
In /root/crol/report.md, write three lines, downtime_requests_failed=, recreate_causes_downtime=, and pdb_min_available=, along with an explanation.
Write three lines, downtime_requests_failed=, recreate_causes_downtime=, and pdb_min_available=, along with an explanation.