TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Break It, Then Fix It

Continue in TT Lab

Goal

You create and fix yourself the five breakages you meet most often in the field, and in doing so you etch into your body the order of going down from symptom to cause.

Why it matters

Troubleshooting is 30% of the CKA score, the largest domain. But you learn nothing by looking only at things that have already been fixed. So this lab is designed in this order: the student creates the broken state, leaves the symptom in a file, and then fixes it. The grading of each step looks not only at the fixed result but also at the evidence that it was broken.

The habit of leaving symptoms is itself a practical skill. When the incident ends, the evidence disappears. To explain the cause later, you have to keep the events and state from that moment.

The cluster state you created in the earlier steps (a cordoned node and a tainted node) remains as it is in the later steps. The last problem can be solved only if you remember that state exactly.

Steps

  1. Create the namespace cka-broken and create the Deployment web (2 replicas) with the image nginx:1.99-nonexistent. Save output showing that wrong image tag to /root/cka-broken/bad-image.txt, then replace the image with nginx:1.27 to reach 2/2 Ready. Do not delete the Deployment; fix only the image.
  2. Create the Pod heavy (image nginx:1.27) with requests cpu: 500 and memory: 2000Gi. Confirm that it becomes Pending and save the output containing the reason to /root/cka-broken/pending.txt. Leave heavy in place without deleting it, and create the Deployment heavy-fixed (1 replica, image nginx:1.27, requests cpu: 500m, memory: 2000Mi) and bring it to 1/1 Ready.
  3. Create the Service api-svc (selector app=api, port 80), and create the Deployment api wrongly, with the Pod label and selector app=api-v2. Save the empty endpoints of api-svc to /root/cka-broken/svc-before.txt, then correct api to the label app=api and bring its 2 replicas to Ready. Leave the Service selector as app=api.
  4. Create the namespace cka-broken-quota and create the ResourceQuota cka-quota with pods: "2" and requests.cpu: "1". When you create the Deployment q-app (4 replicas, image nginx:1.27, requests cpu: 200m), only some of the Pods come up. Save an event showing exceeded quota to /root/cka-broken/quota.txt, then raise the quota's pods to 5 and bring it to 4/4 Ready.
  5. In cka-broken, create the Deployment critical (2 replicas, image nginx:1.27, Pod label app=critical) and the PDB critical-pdb (selector app=critical, minAvailable: 2). Try to drain lab-node-2 and save the blocking message to /root/cka-broken/pdb.txt, then lower minAvailable to 1 and cordon + drain lab-node-2 to empty it completely.
  6. Label lab-node-0 with disktype=ssd and add the taint maintenance=true:NoSchedule. Create the Pod needs-toleration (image nginx:1.27, nodeSelector disktype=ssd, no toleration), confirm that it is Pending, and save the reason to /root/cka-broken/taint.txt. Leaving needs-toleration in place, create the Pod needs-toleration-fixed with the same nodeSelector plus a toleration and place it on lab-node-0.
  7. Create the Pod data-app (image nginx:1.27) so that it mounts the PVC app-data at /data. Save the reason it cannot start because the PVC does not exist yet to /root/cka-broken/pvc.txt, then create the PV cka-fix-pv (1Gi, RWO, storageClassName cka-fix, hostPath /mnt/cka-fix) and the PVC app-data (1Gi, RWO, cka-fix), bind them, and get data-app to Running.
  8. Create the Deployment recovered (4 replicas, image nginx:1.27, Pod label app=recovered) in cka-broken. Requests are cpu: 100m and memory: 128Mi, with a toleration for maintenance=true:NoSchedule, and topologySpreadConstraints of maxSkew 1 / topologyKey kubernetes.io/hostname / whenUnsatisfiable ScheduleAnyway / labelSelector app=recovered. Bring it to 4/4 Ready and summarize the five problems you fixed in /root/cka-broken/summary.md. This file must contain all five words image, taint, selector, quota, and pdb.

Reference

A wrong image tag

Create the namespace cka-broken and create the Deployment web (2 replicas) with the image nginx:1.99-nonexistent. Save output showing that wrong image tag to /root/cka-broken/bad-image.txt, then replace the image with nginx:1.27 to reach 2/2 Ready. Do not delete the Deployment; fix only the image.

First create it with the wrong tag, leave the evidence, and then replace only the image. If you delete the Deployment and create it again, the problem revision disappears and grading catches it.

A resource request unit mistake

Create the Pod heavy (image nginx:1.27) with requests cpu: 500 and memory: 2000Gi. Confirm that it becomes Pending and save the output containing the reason to /root/cka-broken/pending.txt. Leave heavy in place without deleting it, and create the Deployment heavy-fixed (1 replica, image nginx:1.27, requests cpu: 500m, memory: 2000Mi) and bring it to 1/1 Ready.

If you leave out the m in a cpu value, it becomes cores, not millicores. The Events section of describe shows exactly why the scheduler failed.

A Deployment whose labels do not match the Service

Create the Service api-svc (selector app=api, port 80), and create the Deployment api wrongly, with the Pod label and selector app=api-v2. Save the empty endpoints of api-svc to /root/cka-broken/svc-before.txt, then correct api to the label app=api and bring its 2 replicas to Ready. Leave the Service selector as app=api.

A Deployment's spec.selector is immutable. If you created it wrongly, you have to create it again rather than edit it. Do not touch the selector on the Service side.

Exceeding a ResourceQuota

Create the namespace cka-broken-quota and create the ResourceQuota cka-quota with pods: "2" and requests.cpu: "1". When you create the Deployment q-app (4 replicas, image nginx:1.27, requests cpu: 200m), only some of the Pods come up. Save an event showing exceeded quota to /root/cka-broken/quota.txt, then raise the quota's pods to 5 and bring it to 4/4 Ready.

A Pod blocked by the quota shows up not on the Pod but in the ReplicaSet's events. Save that event before you raise the quota.

A drain blocked by a PDB

In cka-broken, create the Deployment critical (2 replicas, image nginx:1.27, Pod label app=critical) and the PDB critical-pdb (selector app=critical, minAvailable: 2). Try to drain lab-node-2 and save the blocking message to /root/cka-broken/pdb.txt, then lower minAvailable to 1 and cordon + drain lab-node-2 to empty it completely.

With 2 replicas and minAvailable 2, you cannot take out even one. Instead of pushing through with --force, recalculate what the PDB is trying to protect.

A missing toleration

Label lab-node-0 with disktype=ssd and add the taint maintenance=true:NoSchedule. Create the Pod needs-toleration (image nginx:1.27, nodeSelector disktype=ssd, no toleration), confirm that it is Pending, and save the reason to /root/cka-broken/taint.txt. Leaving needs-toleration in place, create the Pod needs-toleration-fixed with the same nodeSelector plus a toleration and place it on lab-node-0.

A NoSchedule taint blocks only new Pods and does not touch Pods that are already running. Do not delete the problem Pod; leave it in place.

A PVC that does not exist

Create the Pod data-app (image nginx:1.27) so that it mounts the PVC app-data at /data. Save the reason it cannot start because the PVC does not exist yet to /root/cka-broken/pvc.txt, then create the PV cka-fix-pv (1Gi, RWO, storageClassName cka-fix, hostPath /mnt/cka-fix) and the PVC app-data (1Gi, RWO, cka-fix), bind them, and get data-app to Running.

Create the Pod first and confirm the failure, then create the volume. Once the PVC binds, the scheduler retries that Pod, so you only need to wait a moment.

Putting it together: a recovery deployment on the remaining cluster

Create the Deployment recovered (4 replicas, image nginx:1.27, Pod label app=recovered) in cka-broken. Requests are cpu: 100m and memory: 128Mi, with a toleration for maintenance=true:NoSchedule, and topologySpreadConstraints of maxSkew 1 / topologyKey kubernetes.io/hostname / whenUnsatisfiable ScheduleAnyway / labelSelector app=recovered. Bring it to 4/4 Ready and summarize the five problems you fixed in /root/cka-broken/summary.md. This file must contain all five words image, taint, selector, quota, and pdb.

This cluster now has one cordoned node and one tainted node. Work out what you need to use either one, and then deploy.