CKA — Kubernetes Administrator
Create a Real Failure and Read It
This lab runs on a cluster where real breakage happens
A real k3s is running inside the VM. Because the kubelet and containerd actually run, if an image cannot be pulled you get ImagePullBackOff, if a process dies you get CrashLoopBackOff, and if it exceeds its memory it really gets OOMKilled.
On the fake cluster where the other labs in the CKA course run, you cannot even create these symptoms. So until now you have learned them only from text.
It takes about 2 minutes to start up the first time.
Goal
You deliberately create six kinds of breakage and get hands-on practice with where and how to read each one. At the end you summarize them in a single triage table.
Why it matters
Problem solving is not knowledge but reflex. When you see the words that appear in kubectl get pod, the "next place to look" should come to mind within 3 seconds.
And symptoms that look similar have completely different causes.
PendingandContainerCreatingboth mean "it hasn't come up yet," but the former is a problem of the scheduler and the latter a problem of the kubelet.- A
livenessfailure and areadinessfailure both mean "the probe failed," but the former triggers a restart and the latter only removes the Pod from traffic. - Referencing a nonexistent ConfigMap through
envgivesCreateContainerConfigError, while referencing the same thing as a volume givesPending.
You cannot memorize these distinctions. They stay in your body only if you create them yourself.
Steps
Create all Pods in the broken namespace. The names are fixed.
ts-image— CreateImagePullBackOffwith a nonexistent image, read the actual reason from the events, and put it in/root/cka/imagepull.txt. Then bring up the fixed version asts-image-fixed.ts-crash— CreateCrashLoopBackOffwith a command that dies immediately, read the logs of the container that has already died, and put them in/root/cka/crashloop.txt. Also write down the restart count.ts-oom— Using a container that exceeds its memorylimits, produce anOOMKilled, and put it in/root/cka/oom.txttogether with the exit code.ts-livenessandts-readiness— Set just one failing probe on each, and put how the results differ in/root/cka/probe.txt.ts-pending— Create a Pod that cannot be scheduled and put the reason the scheduler left in/root/cka/pending.txt. Also write down the difference betweenPendingandContainerCreating.ts-config(env reference) andts-volume(volume reference) — Make both point to a nonexistent ConfigMap and put why the statuses differ in/root/cka/configref.txt.- In
/root/cka/triage.md, build a triage table of five symptoms. Each row is증상 · 어디를 볼까 · 흔한 원인(symptom · where to look · common cause). - In
/root/cka/report.md, write four lines,oom_exit_code=,liveness_restarts=,readiness_restarts=, andbroken_pods=, and what you look at first.
Reference
- Events are at the bottom of
kubectl -n broken describe pod <이름>(where the placeholder is the Pod name), underEvents:. You can also see all of them in time order withkubectl get events --sort-by=.lastTimestamp. - The log of a container that has already died is
kubectl logs <파드> --previous(where the placeholder is the Pod). If you do not know this, you will never diagnose CrashLoopBackOff — because the current container has not yet left any log. - The termination reason and code are in
kubectl get pod <이름> -o jsonpath='{.status.containerStatuses[0].lastState.terminated}'(where the placeholder is the Pod name). - To use memory on purpose,
dd if=/dev/zero of=/dev/shm/x bs=1M count=<큰 수>is simple (where the placeholder is a large number). - Common mistake 1: giving only
requestsand nolimitsin step 3. OOM happens when you exceedlimits. - Common mistake 2: putting both probes on one Pod in step 4. Then you cannot tell which probe caused the restart.
When an image cannot be pulled
ts-image — Create ImagePullBackOff with a nonexistent image, read the actual reason from the events, and put it in /root/cka/imagepull.txt. Then bring up the fixed version as ts-image-fixed.
Just specify a tag that does not exist. Do not look only at the status; read the actual reason in the Events of describe.
Read the log of a container that has already died
ts-crash — Create CrashLoopBackOff with a command that dies immediately, read the logs of the container that has already died, and put them in /root/cka/crashloop.txt. Also write down the restart count.
kubectl logs <파드> (where the placeholder is the Pod) looks at the current container. To see the one that just died, you need --previous.
Exceed memory and it dies with 137
ts-oom — Using a container that exceeds its memory limits, produce an OOMKilled, and put it in /root/cka/oom.txt together with the exit code.
Give a small limits.memory and make it use more than that. Also write down what the exit code means.
The two probes have different results
ts-liveness and ts-readiness — Set just one failing probe on each, and put how the results differ in /root/cka/probe.txt.
Put only one probe on a Pod. If you put both, you cannot tell which one caused the restart.
Pending is a scheduler problem
ts-pending — Create a Pod that cannot be scheduled and put the reason the scheduler left in /root/cka/pending.txt. Also write down the difference between Pending and ContainerCreating.
Just request a resource the node cannot give. The scheduler leaves in the events why it could not place the Pod.
Same cause, different symptom
ts-config (env reference) and ts-volume (volume reference) — Make both point to a nonexistent ConfigMap and put why the statuses differ in /root/cka/configref.txt.
Create one Pod that references the nonexistent ConfigMap with envFrom and one that references it as a volume, and compare the statuses.
Summarize in one table
In /root/cka/triage.md, build a triage table of five symptoms. Each row is 증상 · 어디를 볼까 · 흔한 원인 (symptom · where to look · common cause).
Each row is 증상 · 어디를 볼까 · 흔한 원인 (symptom · where to look · common cause). The goal is to make the next action come to mind within 3 seconds in the exam room.
What you learned
In /root/cka/report.md, write four lines, oom_exit_code=, liveness_restarts=, readiness_restarts=, and broken_pods=, and what you look at first.
Write the investigation order together with the four lines oom_exit_code=, liveness_restarts=, readiness_restarts=, and broken_pods=.