CKA — Kubernetes Administrator
Break It, Then Fix It
Goal
You create and fix yourself the five breakages you meet most often in the field, and in doing so you etch into your body the order of going down from symptom to cause.
Why it matters
Troubleshooting is 30% of the CKA score, the largest domain. But you learn nothing by looking only at things that have already been fixed. So this lab is designed in this order: the student creates the broken state, leaves the symptom in a file, and then fixes it. The grading of each step looks not only at the fixed result but also at the evidence that it was broken.
The habit of leaving symptoms is itself a practical skill. When the incident ends, the evidence disappears. To explain the cause later, you have to keep the events and state from that moment.
The cluster state you created in the earlier steps (a cordoned node and a tainted node) remains as it is in the later steps. The last problem can be solved only if you remember that state exactly.
Steps
- Create the namespace
cka-brokenand create the Deploymentweb(2 replicas) with the imagenginx:1.99-nonexistent. Save output showing that wrong image tag to/root/cka-broken/bad-image.txt, then replace the image withnginx:1.27to reach 2/2 Ready. Do not delete the Deployment; fix only the image. - Create the Pod
heavy(imagenginx:1.27) with requestscpu: 500andmemory: 2000Gi. Confirm that it becomes Pending and save the output containing the reason to/root/cka-broken/pending.txt. Leaveheavyin place without deleting it, and create the Deploymentheavy-fixed(1 replica, imagenginx:1.27, requestscpu: 500m,memory: 2000Mi) and bring it to 1/1 Ready. - Create the Service
api-svc(selectorapp=api, port 80), and create the Deploymentapiwrongly, with the Pod label and selectorapp=api-v2. Save the empty endpoints ofapi-svcto/root/cka-broken/svc-before.txt, then correctapito the labelapp=apiand bring its 2 replicas to Ready. Leave the Service selector asapp=api. - Create the namespace
cka-broken-quotaand create the ResourceQuotacka-quotawithpods: "2"andrequests.cpu: "1". When you create the Deploymentq-app(4 replicas, imagenginx:1.27, requestscpu: 200m), only some of the Pods come up. Save an event showingexceeded quotato/root/cka-broken/quota.txt, then raise the quota'spodsto5and bring it to 4/4 Ready. - In
cka-broken, create the Deploymentcritical(2 replicas, imagenginx:1.27, Pod labelapp=critical) and the PDBcritical-pdb(selectorapp=critical,minAvailable: 2). Try to drainlab-node-2and save the blocking message to/root/cka-broken/pdb.txt, then lowerminAvailableto1and cordon + drainlab-node-2to empty it completely. - Label
lab-node-0withdisktype=ssdand add the taintmaintenance=true:NoSchedule. Create the Podneeds-toleration(imagenginx:1.27, nodeSelectordisktype=ssd, no toleration), confirm that it is Pending, and save the reason to/root/cka-broken/taint.txt. Leavingneeds-tolerationin place, create the Podneeds-toleration-fixedwith the same nodeSelector plus a toleration and place it onlab-node-0. - Create the Pod
data-app(imagenginx:1.27) so that it mounts the PVCapp-dataat/data. Save the reason it cannot start because the PVC does not exist yet to/root/cka-broken/pvc.txt, then create the PVcka-fix-pv(1Gi, RWO, storageClassNamecka-fix, hostPath/mnt/cka-fix) and the PVCapp-data(1Gi, RWO,cka-fix), bind them, and getdata-appto Running. - Create the Deployment
recovered(4 replicas, imagenginx:1.27, Pod labelapp=recovered) incka-broken. Requests arecpu: 100mandmemory: 128Mi, with a toleration formaintenance=true:NoSchedule, and topologySpreadConstraints of maxSkew 1 / topologyKeykubernetes.io/hostname/ whenUnsatisfiableScheduleAnyway/ labelSelectorapp=recovered. Bring it to 4/4 Ready and summarize the five problems you fixed in/root/cka-broken/summary.md. This file must contain all five wordsimage,taint,selector,quota, andpdb.
Reference
- The reason scheduling fails is in the Events section of
kubectl describe pod <이름> -n cka-broken(where the placeholder is the Pod name). - A quota overrun shows up not on the Pod but in
kubectl describe rsorkubectl get events -n cka-broken-quota. - A drain goes more smoothly if you add
--ignore-daemonsets --delete-emptydir-data --force --timeout=60s. - Common mistake 1: fixing things before leaving the evidence file. The order is itself a grading item.
- Common mistake 2: leaving out the toleration in step 8.
lab-node-2is cordoned andlab-node-0is tainted, so only one usable node remains.
A wrong image tag
Create the namespace cka-broken and create the Deployment web (2 replicas) with the image nginx:1.99-nonexistent. Save output showing that wrong image tag to /root/cka-broken/bad-image.txt, then replace the image with nginx:1.27 to reach 2/2 Ready. Do not delete the Deployment; fix only the image.
First create it with the wrong tag, leave the evidence, and then replace only the image. If you delete the Deployment and create it again, the problem revision disappears and grading catches it.
A resource request unit mistake
Create the Pod heavy (image nginx:1.27) with requests cpu: 500 and memory: 2000Gi. Confirm that it becomes Pending and save the output containing the reason to /root/cka-broken/pending.txt. Leave heavy in place without deleting it, and create the Deployment heavy-fixed (1 replica, image nginx:1.27, requests cpu: 500m, memory: 2000Mi) and bring it to 1/1 Ready.
If you leave out the m in a cpu value, it becomes cores, not millicores. The Events section of describe shows exactly why the scheduler failed.
A Deployment whose labels do not match the Service
Create the Service api-svc (selector app=api, port 80), and create the Deployment api wrongly, with the Pod label and selector app=api-v2. Save the empty endpoints of api-svc to /root/cka-broken/svc-before.txt, then correct api to the label app=api and bring its 2 replicas to Ready. Leave the Service selector as app=api.
A Deployment's spec.selector is immutable. If you created it wrongly, you have to create it again rather than edit it. Do not touch the selector on the Service side.
Exceeding a ResourceQuota
Create the namespace cka-broken-quota and create the ResourceQuota cka-quota with pods: "2" and requests.cpu: "1". When you create the Deployment q-app (4 replicas, image nginx:1.27, requests cpu: 200m), only some of the Pods come up. Save an event showing exceeded quota to /root/cka-broken/quota.txt, then raise the quota's pods to 5 and bring it to 4/4 Ready.
A Pod blocked by the quota shows up not on the Pod but in the ReplicaSet's events. Save that event before you raise the quota.
A drain blocked by a PDB
In cka-broken, create the Deployment critical (2 replicas, image nginx:1.27, Pod label app=critical) and the PDB critical-pdb (selector app=critical, minAvailable: 2). Try to drain lab-node-2 and save the blocking message to /root/cka-broken/pdb.txt, then lower minAvailable to 1 and cordon + drain lab-node-2 to empty it completely.
With 2 replicas and minAvailable 2, you cannot take out even one. Instead of pushing through with --force, recalculate what the PDB is trying to protect.
A missing toleration
Label lab-node-0 with disktype=ssd and add the taint maintenance=true:NoSchedule. Create the Pod needs-toleration (image nginx:1.27, nodeSelector disktype=ssd, no toleration), confirm that it is Pending, and save the reason to /root/cka-broken/taint.txt. Leaving needs-toleration in place, create the Pod needs-toleration-fixed with the same nodeSelector plus a toleration and place it on lab-node-0.
A NoSchedule taint blocks only new Pods and does not touch Pods that are already running. Do not delete the problem Pod; leave it in place.
A PVC that does not exist
Create the Pod data-app (image nginx:1.27) so that it mounts the PVC app-data at /data. Save the reason it cannot start because the PVC does not exist yet to /root/cka-broken/pvc.txt, then create the PV cka-fix-pv (1Gi, RWO, storageClassName cka-fix, hostPath /mnt/cka-fix) and the PVC app-data (1Gi, RWO, cka-fix), bind them, and get data-app to Running.
Create the Pod first and confirm the failure, then create the volume. Once the PVC binds, the scheduler retries that Pod, so you only need to wait a moment.
Putting it together: a recovery deployment on the remaining cluster
Create the Deployment recovered (4 replicas, image nginx:1.27, Pod label app=recovered) in cka-broken. Requests are cpu: 100m and memory: 128Mi, with a toleration for maintenance=true:NoSchedule, and topologySpreadConstraints of maxSkew 1 / topologyKey kubernetes.io/hostname / whenUnsatisfiable ScheduleAnyway / labelSelector app=recovered. Bring it to 4/4 Ready and summarize the five problems you fixed in /root/cka-broken/summary.md. This file must contain all five words image, taint, selector, quota, and pdb.
This cluster now has one cordoned node and one tainted node. Work out what you need to use either one, and then deploy.