TT Lab
Get started
Learn Learning paths Courses

Kubernetes Distributions — Build Them Yourself

What kubeadm deliberately leaves empty

Continue in TT Lab

One-line summary

kubeadm is not a tool that "installs Kubernetes for you" but a tool that assembles the control plane according to best practices, and the runtime, network, CIDR range, and kernel settings remain the operator's decisions to the very end.

Why this was needed

People who used k3s or k0s first are caught off guard right after kubeadm init succeeds. A green success message appeared, yet the node is NotReady and CoreDNS is Pending. It is not broken. kubeadm deliberately does not pick a network add-on. The sentence the Creating a cluster with kubeadm documentation puts in bold is exactly this: "you must deploy a CNI-based Pod network, and cluster DNS will not start until a network is installed."

The reason for this design is that kubeadm's role is to be "a component of other installation tools." The documentation says it expects kubeadm to be used as a building block in provisioning systems such as Ansible and Terraform, or in larger installation tools. If a component picked the CNI, storage, and Ingress on its own, the tool built on top would have to rip them out again. So kubeadm creates only what must be the same in every environment — certificates, kubeconfigs, static Pod manifests, and bootstrap tokens — and leaves blank whatever varies from environment to environment. The price is that there is more a person has to know.

How it works

The installation is divided into three layers.

호스트 준비      swap · ip_forward · 컨테이너 런타임(CRI) · cgroup 드라이버
kubeadm init    preflight → 인증서(/etc/kubernetes/pki) → kubeconfig → 정적 파드 매니페스트
                → kubelet 이 매니페스트를 보고 etcd·apiserver·controller-manager·scheduler 를 띄움
                → CoreDNS·kube-proxy 애드온, 부트스트랩 토큰, control-plane 테인트
사람의 몫        CNI 설치(대역을 init 과 맞춰서) · 테인트 정리 · 워커 조인 · 인증서 갱신 계획

Host preparation. According to Installing kubeadm, by default the kubelet refuses to start if swap is present. The package repository at pkgs.k8s.io exists separately for each minor version, and the documentation says to exclude kubelet, kubeadm, and kubectl from ordinary upgrades with apt-mark hold after installation, because upgrades must follow the kubeadm procedure. The Container Runtimes documentation explains that the Linux kernel blocks IPv4 forwarding between interfaces by default, and that on a host where systemd is init, using the cgroupfs driver leaves two cgroup managers, which can make the node unstable under resource pressure. On containerd 2.x, SystemdCgroup = true in the runc options under plugins.'io.containerd.cri.v1.runtime' is that switch.

The version matters here. The KubeletCgroupDriverFromCRI feature became stable in 1.34, and if the CRI supports the RuntimeConfig call, the kubelet ignores the cgroupDriver in its own configuration and uses the value the runtime reports. containerd supports this from 2.0. When measured on this lab VM (containerd 2.3.5, kubelet 1.36.4), the kubelet log prints Using cgroup driver setting received from the CRI runtime. In other words, the containerd configuration is now effectively the only source of truth.

init. Preflight mixes checks that block with checks that only warn. Measured on 1.36.4, when ip_forward is 0 it blocks with [ERROR FileContent--proc-sys-net-ipv4-ip_forward], but it does not check br_netfilter. If it passes, kubeadm writes four manifests to /etc/kubernetes/manifests, and the kubelet reads that directory directly to bring up the control plane. This solves the chicken-and-egg problem of having to start the API server when no API server exists yet, using static Pods. The control plane Pods you see in the API server are mirror Pods registered by the kubelet, so their owner is a Node, and even if you delete them the kubelet registers them again.

CIDR ranges. --pod-network-cidr makes the controller-manager hand out a podCIDR to each node, and --service-cidr becomes the API server's --service-cluster-ip-range. The documentation says that problems arise if the Pod network overlaps with the host network, so pick a non-overlapping range and put the same value in both init and the network plugin YAML.

What it looks like in the field

This is a scene reproduced exactly on this lab VM. When you apply the Flannel v0.28.9 manifest, the node turns Ready in 8 seconds. But CoreDNS stays in ContainerCreating for a long time. The reason lies in the order of events. The init container of the Flannel Pod copies /etc/cni/net.d/10-flannel.conflist first, and the kubelet judges the network ready just because the configuration file appeared. Meanwhile the flanneld main process itself dies like this and keeps restarting.

E0914 22:51:43.944688  1 main.go:292] Failed to check br_netfilter:
  stat /proc/sys/net/bridge/bridge-nf-call-iptables: no such file or directory

The Flannel README also says that Flannel needs br_netfilter to start, and that kubeadm has not checked for that module since 1.30. Even after loading the module, I had to wait another 74 seconds (measured) because of the restart interval of CrashLoopBackOff. Ready means "a configuration file exists", not "the data path is working." The judgment must be based on whether CoreDNS actually returns names.

The second scene is about CIDR ranges. This VM runs inside the host Kubernetes, and the VM Pod's address is 10.244.x, the host Pod CIDR. The default Network in the Flannel manifest is exactly 10.244.0.0/16, so if you apply the manifest without editing it, the inner Pod CIDR overlaps the outer one. That is why this lab uses 172.20.0.0/16 · 172.21.0.0/16 and adjusts one line of the manifest to match the init value.

The third is the containerd drop-in. On the author's GPU node (containerd 1.7.27), a single conf.d drop-in once overwrote the entire CRI plugin configuration and the runtime settings went back to their defaults. When I ran the same experiment on this VM's containerd 2.3.5, the drop-in was merged field by field and SystemdCgroup = true remained. Merge rules can differ from version to version, so don't trust a file just by reading it; the answer is the habit of checking with containerd config dump and crictl info.

What really matters in practice

Certificates are valid for one year. According to Certificate Management with kubeadm, client certificates created by kubeadm expire after one year, and the defaults are 8760h for leaf certificates and 87600h for the CA. On this VM too, kubeadm certs check-expiration shows 364 days for leaf certificates and 9 years for the CA. The kubelet certificate is rotated automatically, so it is not in the list. This is where the incident comes from in which a cluster that has gone a whole year without a single upgrade suddenly has every API call rejected one day.

The hash in the join command is a security mechanism. --discovery-token-ca-cert-hash is the sha256 of the CA public key, and it is the value with which the joining node checks "was this API server really signed by our CA?" The token is stored as a bootstrap-token-<id> Secret in kube-system, and the BootstrapSigner attaches a per-token HMAC signature to the cluster-info ConfigMap in kube-public. There is an option to turn off hash verification, but the documentation advises using another method where possible.

If it is a single-machine cluster, remember the taint. For security, kubeadm puts node-role.kubernetes.io/control-plane:NoSchedule on the control plane node so that it does not accept ordinary Pods. If there is only one node, every workload becomes Pending.

Pin the version. If you omit --kubernetes-version, kubeadm first looks up the latest version number on the internet (measured: remote version is much newer: v1.37.0; falling back to: stable-1.36). In an air-gapped network, you get stuck on that one line.

What you will do in the next lab

Starting with the containerd configuration, you set up the control plane with kubeadm init, leave NotReady behind as evidence, and then attach Flannel. You experience for yourself a state where the node is Ready but DNS does not work, and bring it back to life with br_netfilter. Finally, you delete a static Pod, remove the taint, calculate the certificate expiration dates and the join command yourself, and summarize them in a report.