CKAD — Kubernetes Application Developer
Probe Design and Disruption Budgets
Goal
You create the three kinds of probes with the check method (httpGet/tcpSocket/exec) and the timing fields specified, and you design a Deployment's readiness and termination behavior, as well as a PodDisruptionBudget.
Why it matters
A Deployment with no probes attached has all of its rolling update safety net neutralized. Even with maxUnavailable: 0, a container is considered Ready as soon as it starts, and traffic flows into a Pod that is still initializing. Conversely, if you set probes too aggressively, healthy Pods keep restarting and you actually cause an outage. So with probes, the point is not "did you attach them" but "did you attach them with calculated numbers."
If you confuse liveness and readiness, incidents grow. If you check a dependency (the DB) with liveness when it is briefly down, all the Pods restart repeatedly at the same time, and recovery is delayed even after the DB comes back. Check external dependencies with readiness — it is enough for the Pod to stop receiving traffic, and there is no reason to restart.
The startup probe eases the tension between the two. If you give liveness initialDelaySeconds: 300, failure detection during operation is also delayed by 5 minutes, but if you give a separate 300-second budget with startup, you can keep liveness tight after startup.
Steps
- Create the namespace
ckad-obsand create a Podweb-live. Imagenginx:1.27,livenessProbeishttpGetwith path/healthz, port80,initialDelaySeconds: 5,periodSeconds: 10. - Create a Pod
db-ready. Imagenginx:1.27,readinessProbeistcpSocketport5432,initialDelaySeconds: 10,periodSeconds: 5,failureThreshold: 3. - Create a Pod
file-check. Imagebusybox:1.36,command: ["/bin/sh","-c","sleep 3600"],livenessProbeisexecwithcommand: ["cat","/tmp/healthy"],periodSeconds: 5,failureThreshold: 2. - Create a Pod
legacy-app. Imagenginx:1.27.startupProbeishttpGetpath/startup, port8080,failureThreshold: 30,periodSeconds: 10(= a 300-second budget). Also attach to the same container alivenessProbeofhttpGetpath/healthz, port8080,periodSeconds: 10. - Create a Pod
budget-app. Imagenginx:1.27,readinessProbeishttpGetpath/ready, port80,initialDelaySeconds: 15. ChooseperiodSecondsandfailureThresholdyourself so that it is removed from the endpoints within 20 seconds after the first failure. The conditions areperiodSeconds × failureThreshold ≤ 20,failureThreshold ≥ 3,periodSeconds ≥ 2. - Create a Deployment
api. 3 replicas, labelapp=api, imagenginx:1.27. Attach to the container areadinessProbe(httpGet/ready:8080) and alivenessProbe(httpGet/healthz:8080), and specifyterminationMessagePolicy: FallbackToLogsOnError. - Create a PodDisruptionBudget
api-pdb.minAvailable: 2,selectorisapp=api. - Add a
startupProbeto the Deploymentapi—httpGetpath/startup, port8080,failureThreshold: 12,periodSeconds: 5. Then set the Pod spec'sterminationGracePeriodSecondsto60. Keep the existing readiness/liveness probes and the 3 replicas as they are.
Notes
- Check the field names with
kubectl explain pod.spec.containers.livenessProbe/.startupProbe. - It is quicker to extract a skeleton with
kubectl create deployment api --image=nginx:1.27 --replicas=3 -n ckad-obs --dry-run=client -o yaml > api.yamland fill in the probes by hand. - In this environment,
kubectl logsandkubectl execdo not work. Those commands are covered in this module's quiz. - Common mistake 1: writing a probe at the Pod level (
spec.livenessProbe). It is a container-level field. - Common mistake 2: putting the Deployment name in the PDB's
selector. It must be a Pod label. - Common mistake 3: writing
terminationGracePeriodSecondsunder the container. It is at the Pod spec level.
HTTP liveness probe
Create the namespace ckad-obs and create a Pod web-live. Image nginx:1.27, livenessProbe is httpGet with path /healthz, port 80, initialDelaySeconds: 5, periodSeconds: 10.
Give livenessProbe.httpGet a path and a port. The probe is under the container, not at the Pod level. A response code of 200–399 counts as success.
TCP readiness probe
Create a Pod db-ready. Image nginx:1.27, readinessProbe is tcpSocket port 5432, initialDelaySeconds: 10, periodSeconds: 5, failureThreshold: 3.
tcpSocket needs only a port — it succeeds if a connection is established. Even when readiness fails, it does not restart; the Pod is only removed from the endpoints.
exec probe
Create a Pod file-check. Image busybox:1.36, command: ["/bin/sh","-c","sleep 3600"], livenessProbe is exec with command: ["cat","/tmp/healthy"], periodSeconds: 5, failureThreshold: 2.
exec.command is a string array and does not go through a shell. An exit code of 0 means success. To check only that a file exists, there is no need to invoke a shell.
Wrapping a slow startup with a startup probe
Create a Pod legacy-app. Image nginx:1.27. startupProbe is httpGet path /startup, port 8080, failureThreshold: 30, periodSeconds: 10 (= a 300-second budget). Also attach to the same container a livenessProbe of httpGet path /healthz, port 8080, periodSeconds: 10.
Until the startup probe succeeds, liveness/readiness do not start at all. Calculate the startup budget as periodSeconds × failureThreshold. Attach a liveness probe to this Pod as well.
Choosing values by calculating the detection delay
Create a Pod budget-app. Image nginx:1.27, readinessProbe is httpGet path /ready, port 80, initialDelaySeconds: 15. Choose periodSeconds and failureThreshold yourself so that it is removed from the endpoints within 20 seconds after the first failure. The conditions are periodSeconds × failureThreshold ≤ 20, failureThreshold ≥ 3, periodSeconds ≥ 2.
The time from the first failure to action is roughly periodSeconds × failureThreshold. Stay within the required upper bound while setting failureThreshold large enough that a transient failure does not shake it. There is more than one correct answer.
Attaching probes and a termination message policy to a Deployment
Create a Deployment api. 3 replicas, label app=api, image nginx:1.27. Attach to the container a readinessProbe (httpGet /ready:8080) and a livenessProbe (httpGet /healthz:8080), and specify terminationMessagePolicy: FallbackToLogsOnError.
A Deployment's probes go under spec.template.spec.containers[]. The termination message policy is a field at the same container level, and it has only two values, so check with kubectl explain.
Limiting voluntary disruptions with a PodDisruptionBudget
Create a PodDisruptionBudget api-pdb. minAvailable: 2, selector is app=api.
A PDB is policy/v1 and uses only one of minAvailable or maxUnavailable. The selector must point to Pod labels, not the Deployment.
Startup budget and termination budget together (comprehensive)
Add a startupProbe to the Deployment api — httpGet path /startup, port 8080, failureThreshold: 12, periodSeconds: 5. Then set the Pod spec's terminationGracePeriodSeconds to 60. Keep the existing readiness/liveness probes and the 3 replicas as they are.
Add a startup probe to the Deployment you made earlier and lengthen the termination grace period. terminationGracePeriodSeconds is at the Pod spec (spec.template.spec) level, not the container level. Leave the existing liveness/readiness as they are.