CKA — Kubernetes Administrator
Reading Symptoms Is a Reflex
Summary
Problem solving is not learned from text. When you see the words that appear in kubectl get pod, the next place to look should come to mind within 3 seconds, and that stays in your body only if you create the breakage yourself.
Why it has to be real breakage
Problem solving is a large share of the CKA score. But the fake cluster that the earlier modules run on has neither a kubelet nor a container runtime, so it could not create breakage itself.
| Symptom | On the fake cluster |
|---|---|
ImagePullBackOff |
Does not occur, because no image is pulled |
CrashLoopBackOff |
The process does not even die, because there is no process |
OOMKilled |
No memory is used, so it cannot overflow |
| Restart from a probe failure | No probes are run |
PVC Bound |
Stays Pending forever, because there is no provisioner |
So "what does this symptom mean" could be learned only from text. Problem solving is not learned that way. When you see the words that appear in kubectl get pod, the next place to look should come to mind within 3 seconds, and that stays in your body only if you create it yourself.
Things that look similar but have different causes
PendingandContainerCreating— the former is a problem of the scheduler, and the latter of the kubelet. They are told apart by whether the NODE column ofkubectl get pod -o wideis filled in.- A
livenessfailure and areadinessfailure — the former causes a restart, and the latter excludes the Pod from traffic only. - Referencing a nonexistent ConfigMap through
envgivesCreateContainerConfigError, while referencing the same thing as a volume givesPending.
Only three things to remember
- Look at the Events of
describefirst. The status name tells you only where it is stuck. What is missing is always in the events. - If
logsis empty, use--previous. The current container of a CrashLoopBackOff Pod has not printed anything yet. - Look at the NODE column. If it is empty, it is the scheduler; if it is filled in, it is the kubelet. This one line halves your investigation scope.
Where to look next for each symptom
Do not link a status name directly to a cause; link it to where to look next, and the investigation gets much faster. Below is that correspondence table.
| Status | Where to look next | Common causes |
|---|---|---|
Pending, NODE empty |
Events in describe pod |
requests larger than the node's headroom, a taint, a node selector that matches no node |
Pending, volume related |
PVC status and StorageClass | No provisioner, or the access modes do not match |
A long ContainerCreating |
The kubelet on that node | The image is large, a volume mount does not finish, or a Secret is missing |
ImagePullBackOff |
The pull failure line in Events | Tag typo, no credentials for a private registry, network block |
CrashLoopBackOff |
logs --previous |
Exits immediately because of missing configuration, cannot reach a dependent service |
OOMKilled |
Last State in describe |
limits are smaller than actual usage |
CreateContainerConfigError |
The referenced ConfigMap or Secret | A name typo, or not created yet |
No traffic though Running |
Endpoints and readiness | The selector does not match the labels, or readiness keeps failing |
The last row in particular catches people often. The Pod is Running and the logs are normal, but not a single request comes in. What you should look at then is not the Pod but the Service's endpoints. If they are empty, it is one of two things. Either the Service's selector does not match the Pod labels, or the readiness probe fails and the Pod drops out of the list. Telling the two apart is simple. If the Pod shows Ready 0/1, it is a probe problem, and if it shows Ready 1/1 yet the endpoints are empty, it is a selector problem.
When the node itself is the problem, it first shows up as a Pod symptom. When a node becomes NotReady, the Pods on it stay Running for a while and are cleaned up only after a certain time has passed. So in the situation "the Pod is Running but there is no response," you need the habit of running kubectl get nodes once instead of looking only at the Pod. That one line saves tens of minutes.
What really matters in practice
Read the Events, not the status name, first. The status tells you only where it is stuck, and what is missing is always in the events. On the exam and in practice, if you change this order, you spend twice the time.
One NODE column halves the investigation scope. If it is empty, it is the scheduler; if it is filled in, it is the kubelet. If you make kubectl get pod -o wide a habit, this branch comes free.
With CrashLoopBackOff, do not look at logs without --previous. The current container was just recreated and has not printed anything yet. Looking at an empty log and judging "there is no log" is the most common wasted step.
In the next two labs you create these yourself.