TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Ready and Working Are Different Claims

Continue in TT Lab

Summary

Troubleshooting is 30% of the CKA score, the largest of the five domains. And the knack is just one thing: check the state one layer at a time from top to bottom and find the layer where the story breaks off. The Pod status is the starting point of a diagnosis.

Why this matters

If you try to guess the cause from the symptom alone, you will be wrong. The same "I can't connect" can be a selector typo, a scheduling failure, or an unbound PVC. So you fix an order in advance.

Pod status Where it is stuck What to look at first
Pending The scheduler's filter phase Events in describe, resource requests, taints, nodeSelector, PVC binding
ContainerCreating The kubelet's volume/network preparation Volume mounts, CNI, image
ImagePullBackOff Pulling the image Tag typo, registry authentication
CrashLoopBackOff The container keeps dying Logs, configuration, probes
Running but no traffic The Service layer Selector, Ready, port

The one that appears most often in this table is Pending, and within it the most frequent cause is a resource request unit mistake. cpu: 500 is not 500 millicores but 500 cores. You must write 500m. memory: 2000Gi is also likely to be a typo for 2000Mi. The scheduler event says this.

0/3 nodes are available: 3 Insufficient cpu.

How it works

An etcd backup is the last insurance for control plane disaster recovery.

ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%Y%m%d%H%M).db   --endpoints=https://127.0.0.1:2379   --cacert=... --cert=... --key=...
etcdctl snapshot status /backup/etcd-*.db --write-out=table

There are three principles of restoring.

  1. Do not overwrite the existing data-dir. Restore into a new path with --data-dir, and start up with the configuration changed to point to that path.
  2. Bring the apiserver down during the restore. If a live apiserver keeps writing, it will diverge from the restored copy.
  3. Match --initial-cluster and the peer URL (2380) exactly. If you have lost quorum, first restore as a single node and then rejoin the members one at a time.

And one thing you must point out: quorum is for failures, not for mistakes. A wrong deletion is replicated to all members immediately. Only a snapshot allows point-in-time recovery. Also, an etcd snapshot does not contain the data inside PVs. Application data needs a separate means such as Velero.

What it looks like in the field

Case 1 — Everything is Ready but it doesn't work. When I installed KubeVirt on a homelab, the component status was all AllComponentsReady, but the VMs would not start. When I took apart the virt-launcher Pod spec, the volume mount holding the binary that the init container was supposed to run was missing. To quote the author's own words, it was the moment I confirmed, for the third time on that cluster alone, that "a status of Ready and actually working are different propositions." A higher-level status field tells you only what its controller knows about.

Case 2 — Having the binary and having it configured are different. While installing the GPU Operator, I saw that the nvidia-ctk binary was on the node and turned off the toolkit installation. The result was this.

Failed to create pod sandbox: rpc error: code = Unknown
  desc = failed to get sandbox runtime: no runtime for "nvidia" is configured

The Pod could not even start. Nothing failed inside the container; it was blocked at the stage of creating the sandbox. When I checked, kubectl get runtimeclass had nvidia, but on the node grep -c nvidia /etc/containerd/config.toml gave 0. A RuntimeClass is only a name tag like handler: nvidia, and the substance has to be in the node's runtime configuration, but when I rebuilt and switched the runtime from cri-dockerd to containerd, that configuration was lost. The binary on the host was a leftover of the past.

Even after I fixed it, config.toml was still 0. The configuration was in a drop-in file, /etc/containerd/conf.d/99-nvidia.toml. Believing that the one place I grepped was everything was the second trap.

Case 3 — Having quorum doesn't mean you don't need backups. Even after I grew the control plane to 3 machines and secured 3 etcd members, the second line of the remaining task list was this: "Regular etcd snapshots — quorum is for failures, not for mistakes (accidental deletion)." Even with the data replicated three times, a single wrong kubectl delete is reflected in all three copies immediately.

What to do in the next lab

In the first lab you actually take an etcd snapshot, compare resources before and after the backup to see with your own eyes what the snapshot does not contain, and leave a recovery plan as a document. In the second lab you create and fix five kinds of breakage yourself. A wrong image tag, a resource request unit mistake, a missing toleration, a label mismatch, exceeding a ResourceQuota, and a drain blocked by a PDB. What is graded is leaving the symptom in a file before you fix it.