TT Lab
Get started
Learn Learning paths Courses

CKAD — Kubernetes Application Developer

Probe Design and Disruption Budgets

Continue in TT Lab

Goal

You create the three kinds of probes with the check method (httpGet/tcpSocket/exec) and the timing fields specified, and you design a Deployment's readiness and termination behavior, as well as a PodDisruptionBudget.

Why it matters

A Deployment with no probes attached has all of its rolling update safety net neutralized. Even with maxUnavailable: 0, a container is considered Ready as soon as it starts, and traffic flows into a Pod that is still initializing. Conversely, if you set probes too aggressively, healthy Pods keep restarting and you actually cause an outage. So with probes, the point is not "did you attach them" but "did you attach them with calculated numbers."

If you confuse liveness and readiness, incidents grow. If you check a dependency (the DB) with liveness when it is briefly down, all the Pods restart repeatedly at the same time, and recovery is delayed even after the DB comes back. Check external dependencies with readiness — it is enough for the Pod to stop receiving traffic, and there is no reason to restart.

The startup probe eases the tension between the two. If you give liveness initialDelaySeconds: 300, failure detection during operation is also delayed by 5 minutes, but if you give a separate 300-second budget with startup, you can keep liveness tight after startup.

Steps

  1. Create the namespace ckad-obs and create a Pod web-live. Image nginx:1.27, livenessProbe is httpGet with path /healthz, port 80, initialDelaySeconds: 5, periodSeconds: 10.
  2. Create a Pod db-ready. Image nginx:1.27, readinessProbe is tcpSocket port 5432, initialDelaySeconds: 10, periodSeconds: 5, failureThreshold: 3.
  3. Create a Pod file-check. Image busybox:1.36, command: ["/bin/sh","-c","sleep 3600"], livenessProbe is exec with command: ["cat","/tmp/healthy"], periodSeconds: 5, failureThreshold: 2.
  4. Create a Pod legacy-app. Image nginx:1.27. startupProbe is httpGet path /startup, port 8080, failureThreshold: 30, periodSeconds: 10 (= a 300-second budget). Also attach to the same container a livenessProbe of httpGet path /healthz, port 8080, periodSeconds: 10.
  5. Create a Pod budget-app. Image nginx:1.27, readinessProbe is httpGet path /ready, port 80, initialDelaySeconds: 15. Choose periodSeconds and failureThreshold yourself so that it is removed from the endpoints within 20 seconds after the first failure. The conditions are periodSeconds × failureThreshold ≤ 20, failureThreshold ≥ 3, periodSeconds ≥ 2.
  6. Create a Deployment api. 3 replicas, label app=api, image nginx:1.27. Attach to the container a readinessProbe (httpGet /ready:8080) and a livenessProbe (httpGet /healthz:8080), and specify terminationMessagePolicy: FallbackToLogsOnError.
  7. Create a PodDisruptionBudget api-pdb. minAvailable: 2, selector is app=api.
  8. Add a startupProbe to the Deployment api — httpGet path /startup, port 8080, failureThreshold: 12, periodSeconds: 5. Then set the Pod spec's terminationGracePeriodSeconds to 60. Keep the existing readiness/liveness probes and the 3 replicas as they are.

Notes

HTTP liveness probe

Create the namespace ckad-obs and create a Pod web-live. Image nginx:1.27, livenessProbe is httpGet with path /healthz, port 80, initialDelaySeconds: 5, periodSeconds: 10.

Give livenessProbe.httpGet a path and a port. The probe is under the container, not at the Pod level. A response code of 200–399 counts as success.

TCP readiness probe

Create a Pod db-ready. Image nginx:1.27, readinessProbe is tcpSocket port 5432, initialDelaySeconds: 10, periodSeconds: 5, failureThreshold: 3.

tcpSocket needs only a port — it succeeds if a connection is established. Even when readiness fails, it does not restart; the Pod is only removed from the endpoints.

exec probe

Create a Pod file-check. Image busybox:1.36, command: ["/bin/sh","-c","sleep 3600"], livenessProbe is exec with command: ["cat","/tmp/healthy"], periodSeconds: 5, failureThreshold: 2.

exec.command is a string array and does not go through a shell. An exit code of 0 means success. To check only that a file exists, there is no need to invoke a shell.

Wrapping a slow startup with a startup probe

Create a Pod legacy-app. Image nginx:1.27. startupProbe is httpGet path /startup, port 8080, failureThreshold: 30, periodSeconds: 10 (= a 300-second budget). Also attach to the same container a livenessProbe of httpGet path /healthz, port 8080, periodSeconds: 10.

Until the startup probe succeeds, liveness/readiness do not start at all. Calculate the startup budget as periodSeconds × failureThreshold. Attach a liveness probe to this Pod as well.

Choosing values by calculating the detection delay

Create a Pod budget-app. Image nginx:1.27, readinessProbe is httpGet path /ready, port 80, initialDelaySeconds: 15. Choose periodSeconds and failureThreshold yourself so that it is removed from the endpoints within 20 seconds after the first failure. The conditions are periodSeconds × failureThreshold ≤ 20, failureThreshold ≥ 3, periodSeconds ≥ 2.

The time from the first failure to action is roughly periodSeconds × failureThreshold. Stay within the required upper bound while setting failureThreshold large enough that a transient failure does not shake it. There is more than one correct answer.

Attaching probes and a termination message policy to a Deployment

Create a Deployment api. 3 replicas, label app=api, image nginx:1.27. Attach to the container a readinessProbe (httpGet /ready:8080) and a livenessProbe (httpGet /healthz:8080), and specify terminationMessagePolicy: FallbackToLogsOnError.

A Deployment's probes go under spec.template.spec.containers[]. The termination message policy is a field at the same container level, and it has only two values, so check with kubectl explain.

Limiting voluntary disruptions with a PodDisruptionBudget

Create a PodDisruptionBudget api-pdb. minAvailable: 2, selector is app=api.

A PDB is policy/v1 and uses only one of minAvailable or maxUnavailable. The selector must point to Pod labels, not the Deployment.

Startup budget and termination budget together (comprehensive)

Add a startupProbe to the Deployment api — httpGet path /startup, port 8080, failureThreshold: 12, periodSeconds: 5. Then set the Pod spec's terminationGracePeriodSeconds to 60. Keep the existing readiness/liveness probes and the 3 replicas as they are.

Add a startup probe to the Deployment you made earlier and lengthen the termination grace period. terminationGracePeriodSeconds is at the Pod spec (spec.template.spec) level, not the container level. Leave the existing liveness/readiness as they are.