TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

What kubeadm Does Not Do for You — Preparation, Skew, Interfaces

Continue in TT Lab

Summary

kubeadm is a tool that assembles the control plane, not a tool that prepares the nodes. Swap, IP forwarding, the container runtime, the cgroup driver, and ports have to be set up by a person; an upgrade has to follow the order the version skew policy sets (kube-apiserver first); and CNI, CSI, and CRI are interfaces that split networking, storage, and the runtime out of the core so they can be swapped.

Why this matters

kubeadm init runs preflight checks as soon as it starts. The Installing kubeadm document states the requirements like this — at least 2GB of RAM per machine, at least 2 CPUs for a control plane machine, full network connectivity between all machines, a unique hostname, MAC address, and product_uuid on every node, and certain ports being open. The binaries are dynamically linked to glibc, so a distribution without glibc such as Alpine needs a compatibility layer, and kubeadm checks whether the kernel version is supported with the SystemVerification check.

When these checks fail, kubeadm stops, but it does not fix things so that they pass. So if you do not know the preparation items, you start from "why isn't this working" and fill in the gaps one by one by searching, and in the exam room you do not have that time. The same goes for upgrades. kubeadm enforces the skew policy, but a person has to know which nodes to upgrade and in which order.

How it works

Infrastructure preparation before installation

Swap. The kubelet's default behavior is to refuse to start if it detects swap on the node. So you turn it off with swapoff -a and remove it from /etc/fstab or the systemd.swap configuration so that it stays off after a reboot. If you really want to keep swap, you have to give failSwapOn: false in the kubelet configuration, and even then workloads do not use swap because of the default swapBehavior of NoSwap.

Kernel network settings. According to the Container Runtimes document, the Linux kernel by default does not allow routing IPv4 packets between interfaces. So, in /etc/sysctl.d/k8s.conf, you write net.ipv4.ip_forward = 1 and apply it with sysctl --system. The document says most network implementations change this value themselves when needed but some leave it to the administrator, and adds that some implementations expect other sysctl values or kernel module loading, so you should read the network implementation's documentation. Loading the overlay and br_netfilter modules and the bridge-nf-call family of sysctls, which were in older guides, are not on the current official page's required list and are decided by whether the CNI implementation requires them. Rather than memorizing the old procedure as is for the exam or the field, it is safer to keep apart what the current document requires (ip_forward) and what the implementation requires.

Container runtime. Kubernetes talks to the runtime through CRI. If you do not specify a runtime, kubeadm scans known socket paths and detects one automatically, and if there are several or none it raises an error and asks you to specify one. Docker Engine does not implement CRI, so you must install cri-dockerd separately, and the kubelet's built-in Docker support (dockershim) was removed in 1.24.

Runtime Unix socket
containerd unix:///var/run/containerd/containerd.sock
CRI-O unix:///var/run/crio/crio.sock
Docker Engine (cri-dockerd) unix:///var/run/cri-dockerd.sock

The cgroup driver. The kubelet and the runtime both limit resources with cgroups, and the drivers used for that are the two, cgroupfs and systemd. The rule the document calls "critical" is that the kubelet and the runtime must use the same driver. The kubelet's default is cgroupfs, but on distributions whose init system is systemd, there would be two cgroup managers and the node can become unstable under resource pressure, so you use the systemd driver. If you use cgroup v2, the systemd driver is the answer. On the kubelet side, use the KubeletConfiguration field cgroupDriver: systemd, and on the runtime side it is the setting in each runtime's documentation. In Kubernetes 1.37, if the KubeletCgroupDriverFromCRI feature gate is on and the runtime supports the RuntimeConfig CRI RPC, the kubelet finds out the driver from the runtime automatically, but for a runtime that lacks that RPC, such as containerd 1.y, it still uses the kubelet's own setting.

Ports. Here is the table from the Ports and Protocols document, carried over as is.

Where Port Purpose
Control plane 6443 kube-apiserver
Control plane 2379-2380 etcd client API
Control plane 10250 kubelet API
Control plane 10259 kube-scheduler
Control plane 10257 kube-controller-manager
Worker 10250 kubelet API
Worker 10256 kube-proxy
Worker 30000-32767 NodePort Service (TCP and UDP)

Cluster lifecycle — skew sets the order

The Version Skew Policy says the project maintains the three most recent minor release branches, and that versions after 1.19 get about a year of patch support. The allowed differences between components are as follows.

The upgrade order comes from these rules. You upgrade kube-apiserver first (you cannot skip a minor), then controller-manager and scheduler, and finally the kubelet and kube-proxy. If you do it the other way around, you pass through the forbidden state where "the kubelet is newer than the apiserver." The document recommends upgrading to the latest patch of the current minor before upgrading, and going to the latest patch of the target minor.

The procedure set by the Upgrading kubeadm clusters document has three stages — the first control plane node, additional control plane nodes, and worker nodes.

  1. On the first control plane, check the versions you can upgrade to and the state of the component configuration with kubeadm upgrade plan, then run kubeadm upgrade apply v1.X.y. This command checks whether the API server is reachable, that all nodes are Ready, and the control plane state, enforces the skew policy, prepares the images, replaces the control plane components but rolls back if any one fails to come up, and applies the new CoreDNS and kube-proxy manifests. The certificates that kubeadm manages are also renewed at this point.
  2. On the additional control plane nodes, you use kubeadm upgrade node. It fetches the cluster's ClusterConfiguration and upgrades the static Pod manifests and the kubelet configuration.
  3. When you upgrade a node's kubelet by a minor version, you drain it first. Then you upgrade the kubelet and kubectl packages, on a worker you refresh the kubelet configuration with kubeadm upgrade node and restart the kubelet, and you put it back with kubectl uncordon.

The document tells you in advance that all containers are restarted after an upgrade because the container spec hash changes. On failure you may rerun the same command (it is idempotent), you can also bring the state in line without changing the version with kubeadm upgrade apply --force, and backups of etcd and the manifests remain under /etc/kubernetes/tmp.

According to the Safely Drain a Node document, kubectl drain terminates Pods gracefully and respects PodDisruptionBudgets, and if there are DaemonSet Pods you need --ignore-daemonsets. That a drain finished successfully means that all Pods (except excluded system Pods) were safely evicted, and after that you reopen scheduling with kubectl uncordon. You run it one node at a time, but even if you run it in parallel from several terminals, the PDB is respected together.

CNI, CSI, CRI — what was split out

The three interfaces have similar names but split out different layers.

Interface What was split out Who calls whom
CRI The container runtime The kubelet calls the runtime as a gRPC client
CNI The Pod network implementation The container runtime loads the CNI plugin to implement the Pod network model
CSI The storage system A CSI driver exposes storage through a standard interface, and Pods use it with the csi volume type

The CRI document says CRI is the main gRPC protocol between the kubelet and the runtime, that it became stable in v1.23, and that since 1.26 the kubelet requires the v1 CRI API, so on a runtime that does not support it the node does not register. You specify the endpoint with the kubelet's --container-runtime-endpoint. The Network Plugins document explains that a CNI plugin is required to implement the Kubernetes network model, that it must be compatible with CNI spec v0.4.0 or later (v1.0.0 recommended), and that in 1.24 the kubelet's cni-bin-dir and network-plugin flags were removed, so CNI management became the runtime's job, not the kubelet's. The CSI section of the Volumes document says CSI is a standard interface that exposes arbitrary storage systems to container workloads, that it can be used in three ways (PVC reference, generic ephemeral volume, and CSI ephemeral volume), and that the older approach, FlexVolume, has been deprecated since 1.23.

What it looks like in the field

The node wobbles only under resource pressure. For a node that is fine normally but where the kubelet restarts or Pods die strangely when memory fills up, suspect a cgroup driver mismatch. The document describes exactly this symptom — if systemd is the init but the kubelet and runtime use cgroupfs, there are two cgroup managers and the resource views diverge. The document warns that changing the driver on a node that has already joined the cluster is also a delicate operation, so it is far cheaper to match them before joining.

I upgraded the kubelet without a drain. If you only swap the package and restart the kubelet, every container on that node restarts, and since you did not drain, neither the PDB nor graceful termination is respected. This is why the document states firmly that for a minor upgrade you drain first.

What to check in the next quiz

The quiz asks about the kubelet's default behavior with swap, why you match the cgroup driver to systemd, the kernel settings the current official documentation requires, the upgrade order, the kubelet's allowed skew, and who is the one that loads CNI plugins.