TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Create a Real Failure and Read It

Continue in TT Lab

This lab runs on a cluster where real breakage happens

A real k3s is running inside the VM. Because the kubelet and containerd actually run, if an image cannot be pulled you get ImagePullBackOff, if a process dies you get CrashLoopBackOff, and if it exceeds its memory it really gets OOMKilled.

On the fake cluster where the other labs in the CKA course run, you cannot even create these symptoms. So until now you have learned them only from text.

It takes about 2 minutes to start up the first time.

Goal

You deliberately create six kinds of breakage and get hands-on practice with where and how to read each one. At the end you summarize them in a single triage table.

Why it matters

Problem solving is not knowledge but reflex. When you see the words that appear in kubectl get pod, the "next place to look" should come to mind within 3 seconds.

And symptoms that look similar have completely different causes.

You cannot memorize these distinctions. They stay in your body only if you create them yourself.

Steps

Create all Pods in the broken namespace. The names are fixed.

  1. ts-image — Create ImagePullBackOff with a nonexistent image, read the actual reason from the events, and put it in /root/cka/imagepull.txt. Then bring up the fixed version as ts-image-fixed.
  2. ts-crash — Create CrashLoopBackOff with a command that dies immediately, read the logs of the container that has already died, and put them in /root/cka/crashloop.txt. Also write down the restart count.
  3. ts-oom — Using a container that exceeds its memory limits, produce an OOMKilled, and put it in /root/cka/oom.txt together with the exit code.
  4. ts-liveness and ts-readiness — Set just one failing probe on each, and put how the results differ in /root/cka/probe.txt.
  5. ts-pending — Create a Pod that cannot be scheduled and put the reason the scheduler left in /root/cka/pending.txt. Also write down the difference between Pending and ContainerCreating.
  6. ts-config (env reference) and ts-volume (volume reference) — Make both point to a nonexistent ConfigMap and put why the statuses differ in /root/cka/configref.txt.
  7. In /root/cka/triage.md, build a triage table of five symptoms. Each row is 증상 · 어디를 볼까 · 흔한 원인 (symptom · where to look · common cause).
  8. In /root/cka/report.md, write four lines, oom_exit_code=, liveness_restarts=, readiness_restarts=, and broken_pods=, and what you look at first.

Reference

When an image cannot be pulled

ts-image — Create ImagePullBackOff with a nonexistent image, read the actual reason from the events, and put it in /root/cka/imagepull.txt. Then bring up the fixed version as ts-image-fixed.

Just specify a tag that does not exist. Do not look only at the status; read the actual reason in the Events of describe.

Read the log of a container that has already died

ts-crash — Create CrashLoopBackOff with a command that dies immediately, read the logs of the container that has already died, and put them in /root/cka/crashloop.txt. Also write down the restart count.

kubectl logs <파드> (where the placeholder is the Pod) looks at the current container. To see the one that just died, you need --previous.

Exceed memory and it dies with 137

ts-oom — Using a container that exceeds its memory limits, produce an OOMKilled, and put it in /root/cka/oom.txt together with the exit code.

Give a small limits.memory and make it use more than that. Also write down what the exit code means.

The two probes have different results

ts-liveness and ts-readiness — Set just one failing probe on each, and put how the results differ in /root/cka/probe.txt.

Put only one probe on a Pod. If you put both, you cannot tell which one caused the restart.

Pending is a scheduler problem

ts-pending — Create a Pod that cannot be scheduled and put the reason the scheduler left in /root/cka/pending.txt. Also write down the difference between Pending and ContainerCreating.

Just request a resource the node cannot give. The scheduler leaves in the events why it could not place the Pod.

Same cause, different symptom

ts-config (env reference) and ts-volume (volume reference) — Make both point to a nonexistent ConfigMap and put why the statuses differ in /root/cka/configref.txt.

Create one Pod that references the nonexistent ConfigMap with envFrom and one that references it as a volume, and compare the statuses.

Summarize in one table

In /root/cka/triage.md, build a triage table of five symptoms. Each row is 증상 · 어디를 볼까 · 흔한 원인 (symptom · where to look · common cause).

Each row is 증상 · 어디를 볼까 · 흔한 원인 (symptom · where to look · common cause). The goal is to make the next action come to mind within 3 seconds in the exam room.

What you learned

In /root/cka/report.md, write four lines, oom_exit_code=, liveness_restarts=, readiness_restarts=, and broken_pods=, and what you look at first.

Write the investigation order together with the four lines oom_exit_code=, liveness_restarts=, readiness_restarts=, and broken_pods=.