KCNA — Kubernetes and Cloud Native Associate
Why CRI Exists and What the shim Is Holding On To
One-line summary
The CRI is "the standard grammar in which Kubernetes talks to the runtime," and the shim is "a thin process that holds on to the container so it stays alive even if containerd dies." Both are products of conventions created to get rid of adapters.
Why this was needed
Back when there was no CRI, the only runtime Kubernetes could use was Docker, and Kubernetes carried an adapter inside its own code to talk to the Docker Engine. That is dockershim.
This code was right in its day. Without it, early adoption would have been impossible. The problem shows up as runtimes multiply. If you have to put one more adapter into the body of Kubernetes every time a runtime is added, the maintenance burden piles up on the Kubernetes side without limit.
So the order was flipped. To remove the adapters, you first have to create a convention. A gRPC interface called the CRI (Container Runtime Interface) was defined, and the runtime side was made to implement that interface. dockershim was removed in Kubernetes v1.24, and the official FAQ states that this code was intended from the start as a temporary solution. containerd and CRI-O filled its place.
Here we have to cut off one misconception that comes up often on the exam. The removal of dockershim does not mean "Docker images no longer work." The image format is the OCI Image Spec, so an image built with docker build runs as is on every CRI implementation. What disappeared is not the image but the adapter code inside kubelet.
How it works
The two services the CRI provides
- RuntimeService: The lifecycle of Pod sandboxes and containers.
RunPodSandbox,StopPodSandbox,CreateContainer,StartContainer,StopContainer,ListContainers,ContainerStatus,ExecSync,Exec,Attach,PortForward - ImageService: Images.
PullImage,ListImages,ImageStatus,RemoveImage,ImageFsInfo
What actually happens, in order, when a Pod starts
- kubelet calls
RunPodSandbox - containerd creates the pause container, which is the owner of the Pod's network namespace
- The CNI plugin is called out and attaches an IP to that namespace
- kubelet calls
PullImage→CreateContainer→StartContainer - containerd runs the container through the shim
Why the pause container exists is a KCNA favorite. For the containers in a Pod to share an IP and port space, someone has to create that network namespace first and hold on to it to the end. The namespace must persist even at the moment all the app containers disappear because of restarts, or the IP would change. Because pause just waits forever with the pause() system call, it uses only about 1 MB of memory, and as the init of the PID namespace it also reaps zombie processes.
The shim: a device that decouples containerd restarts from container lifetime
containerd --> containerd-shim-runc-v2 --> runc --> 컨테이너 프로세스
The shim has four responsibilities.
- Keep the container alive even if containerd restarts
- Manage the container's stdin/stdout/stderr
- Collect the exit code
- Report OOM events
The parent of the container process is not containerd but the shim, and the shim daemonizes itself and runs in a different session from containerd. That is why the PID of the container process does not change even if you restart containerd. Restarting containerd because "the container is acting strange" mostly fixes nothing: only the management plane comes up fresh, because what actually holds on to the container is the shim.
The shim is swapped according to the kind of container: containerd-shim-runc-v2 (runc), containerd-shim-kata-v2 (lightweight VM), containerd-shim-runsc-v1 (gVisor), and containerd-shim-wasm (WebAssembly). The RuntimeClass object exposes this choice on a per-Pod basis.
What it looks like in practice
While rebuilding the cluster, the author's homelab changed the CNI from Calico to Cilium 1.20.1 and did not install kube-proxy at all (--skip-phases=addon/kube-proxy). And it confirmed that this actually works not by assertion but by measurement: 0 kube-proxy Pods, and 0 KUBE- chains of iptables on the nodes. Instead, the eBPF map directly held a mapping such as 10.96.0.1:443/TCP → 10.0.0.120:6443/TCP.
Here the talk about layers becomes real. Changing the CNI and changing the CRI are work at different layers. The CNI is in charge of only step 3 (attaching an IP to the namespace), and steps 1, 4, and 5, which create the container, stay the same. So even if you replace the CNI entirely, the conversation between kubelet and containerd does not change by a single character.
One more thing. Thanks to this layer separation, there is a failure where only exec does not work. exec/attach/port-forward do not flow over the CRI gRPC; containerd returns a streaming URL and the client makes a new connection to the node at that URL. So you can get a situation where the Pod is running fine and only kubectl exec times out.
What to check in the next quiz
This module ends with a quiz. From the next module on, you connect to a real cluster and see for yourself how the pieces we have talked about so far show up as API objects.