Kubernetes Distributions — Build Them Yourself
init succeeded, but the node is NotReady
Goal
On an empty Ubuntu VM, set up a single control plane with kubeadm, confirm with evidence that the cause of NotReady is the CNI, and then attach Flannel to bring DNS to life as well. Along the way, you see by hand what each of the static Pod, the taint, the certificate lifetime, and the join token does.
Why it matters
With k3s and k0s, the distribution picks the runtime, CNI, and storage for you. kubeadm is a standard assembly kit that leaves all those decisions to a person, so "init succeeded" does not mean "the cluster is usable." Choosing the network add-on, matching the kernel settings that add-on requires, and making sure the CIDR ranges do not overlap with the surrounding networks are all the operator's job. In this lab you experience for yourself that the job is not done just because the node has turned Ready — the node becomes Ready as soon as a CNI configuration file appears, but if the daemon is dead, CoreDNS does not come up. The one-year certificate expiration and the CA hash of the join token cause no problem on the day of installation and then become an incident a few months later, so you learn to read them ahead of time.
Steps
- Write
net.ipv4.ip_forward = 1to/etc/sysctl.d/k8s.confand apply it, create/etc/containerd/config.tomlwithcontainerd config default, change runc'sSystemdCgroupto true, and restart containerd. Then write to/root/kubeadm/runtime.jsonthe fieldscontainerd_version(the third column ofcontainerd --version),systemd_cgroup(the value the running containerd reports, a boolean),ip_forward(the current kernel value, a number), andswap_total_kb(SwapTotal in /proc/meminfo, a number). - Run
kubeadm initwith--kubernetes-version v1.36.4 --pod-network-cidr 172.20.0.0/16 --service-cidr 172.21.0.0/16and save the entire output to/root/kubeadm/init.log. Copy/etc/kubernetes/admin.confto/root/.kube/configso thatkubectlsees the new cluster. - Before installing the CNI, record the current state in
/root/kubeadm/notready.json. Fields:node_uid,ready_status(the status of the Ready condition),ready_message(the message of the Ready condition),taints(an array holding the node taints as키:효과strings, that is, key:effect),coredns_phase(the phase of one CoreDNS Pod),cni_conf_count(the number of files in/etc/cni/net.dwhose names do not start with a dot),observed_at(the UTC time of recording, in the2026-01-01T00:00:00Zformat). - Download the
kube-flannel.ymlof the Flannel v0.28.9 release to/root/kubeadm/kube-flannel.yml, changeNetworkinnet-conf.jsonto this cluster's Pod CIDR range, and apply it (do not load br_netfilter yet). After watching for about 30 seconds, write to/root/kubeadm/cni-symptom.jsonthe fieldsnode_ready(the Ready condition status),coredns_ready_replicas(readyReplicas of the coredns Deployment, 0 if absent),flannel_pod_uid,flannel_restarts(restartCount of the kube-flannel container), andflannel_error(the one line in the flannel log that states the cause). - Write the
br_netfiltermodule to/etc/modules-load.d/k8s.confso that it loads at boot, and load it now as well. Addnet.bridge.bridge-nf-call-iptables = 1to/etc/sysctl.d/k8s.confand apply it. The Flannel DaemonSet must become ready, both CoreDNS Pods must become Available, and on the node,dig @<kube-dns 서비스 IP> kubernetes.default.svc.cluster.local(the placeholder is the kube-dns Service IP) must return the kubernetes Service IP. - Write to
/root/kubeadm/static-pods.jsonthe fieldsmanifests(a sorted array of the file names in/etc/kubernetes/manifests),owner_kind(the ownerReferences kind of the kube-scheduler Pod),etcd_data_dir(the data path the etcd manifest mounts as a hostPath), andservice_cluster_ip_range(the value of the flag of the same name in the kube-apiserver manifest). Then delete the kube-scheduler Pod withkubectl delete, confirm that it comes back, and adddeleted_uid,new_uid,container_id_before, andcontainer_id_after(containerStatuses[0].containerID) to the same file. - In the default namespace, create a Pod
probe(imageregistry.k8s.io/pause:3.10.2, restartPolicy Never), confirm that it is Pending, and write to/root/kubeadm/pending.jsonthe fieldspod_uid,phase,message(the message of the PodScheduled condition), andobserved_at(UTC, in the2026-01-01T00:00:00Zformat). Then remove thenode-role.kubernetes.io/control-plane:NoScheduletaint from the node so thatprobebecomes Running. Do not delete and recreate the Pod. - Check with
kubeadm certs check-expirationand openssl, and write to/root/kubeadm/certs.jsonthe fieldsapiserver_not_afterandca_not_after(both UTC, in the2026-01-01T00:00:00Zformat), andapiserver_valid_daysandca_valid_days(the number of days from notBefore to notAfter of each certificate, an integer). Then create a new token withkubeadm token create --ttl 3h --description worker-join, calculate the CA public key hash yourself, and write to/root/kubeadm/join.txta single linekubeadm join <API 엔드포인트> --token <토큰> --discovery-token-ca-cert-hash sha256:<해시>(the placeholders are the API endpoint, the token, and the hash). - Write to
/root/kubeadm/report.jsonthe fieldskubernetes_version(the server gitVersion),pod_cidr,service_cidr,dns_service_ip(the kube-dns Service IP),cgroup_driver(the driver the kubelet uses),notready_cause(the cause of the NotReady in step 3: one ofcni·runtime·certificate),flannel_blocker(the name of the kernel module that blocked flannel in step 4), andjoin_token_id(the first 6 characters of the token in step 8).
Notes
- The VM has containerd 2.3.5 (
/usr/local), runc 1.5.1, crictl 1.36.0, and kubeadm, kubelet, and kubectl 1.36.4 installed, and the images are already pulled. There is no cluster yet. - To start over after running init once, you need
kubeadm reset -f. It is hard to undo, so check the CIDR ranges first. - Common mistake: loading br_netfilter before step 4. Then the symptom you are meant to see in step 4 does not appear, so you cannot record it.
- Common mistake: looking only at the result of
containerd config dumpand forgetting to restart. The running value is incrictl info. - Documentation: Installing kubeadm · Creating a cluster with kubeadm · Container Runtimes · Certificate Management with kubeadm
Set up the runtime before init
Write net.ipv4.ip_forward = 1 to /etc/sysctl.d/k8s.conf and apply it, create /etc/containerd/config.toml with containerd config default, change runc's SystemdCgroup to true, and restart containerd. Then write to /root/kubeadm/runtime.json the fields containerd_version (the third column of containerd --version), systemd_cgroup (the value the running containerd reports, a boolean), ip_forward (the current kernel value, a number), and swap_total_kb (SwapTotal in /proc/meminfo, a number).
What you write in a file and what the running daemon uses are different. containerd config dump only shows the merged configuration files, and to find out what a daemon that has not been restarted is using, you have to ask the CRI (crictl info). The current value of ip_forward changes only if you make the system reread the files with sysctl --system.
Run init while avoiding the CIDR ranges
Run kubeadm init with --kubernetes-version v1.36.4 --pod-network-cidr 172.20.0.0/16 --service-cidr 172.21.0.0/16 and save the entire output to /root/kubeadm/init.log. Copy /etc/kubernetes/admin.conf to /root/.kube/config so that kubectl sees the new cluster.
This VM runs inside the host Kubernetes, and the host uses 10.244.0.0/16 for Pods and 10.96.0.0/12 for Services. If the two ranges overlap, the inner kube-proxy intercepts the outer DNS address. If you omit --kubernetes-version, kubeadm goes to the internet to look up the latest version number first.
init succeeded but the node is NotReady
Before installing the CNI, record the current state in /root/kubeadm/notready.json. Fields: node_uid, ready_status (the status of the Ready condition), ready_message (the message of the Ready condition), taints (an array holding the node taints as 키:효과 strings, that is, key:effect), coredns_phase (the phase of one CoreDNS Pod), cni_conf_count (the number of files in /etc/cni/net.d whose names do not start with a dot), observed_at (the UTC time of recording, in the 2026-01-01T00:00:00Z format).
The message of the node's Ready condition states the cause directly. The reason CoreDNS is Pending is in the Pod's PodScheduled condition; check which taint on the node that reason refers to. The grader compares timestamps to check whether this record was written before the CNI installation.
After installing the CNI, only the node became Ready
Download the kube-flannel.yml of the Flannel v0.28.9 release to /root/kubeadm/kube-flannel.yml, change Network in net-conf.json to this cluster's Pod CIDR range, and apply it (do not load br_netfilter yet). After watching for about 30 seconds, write to /root/kubeadm/cni-symptom.json the fields node_ready (the Ready condition status), coredns_ready_replicas (readyReplicas of the coredns Deployment, 0 if absent), flannel_pod_uid, flannel_restarts (restartCount of the kube-flannel container), and flannel_error (the one line in the flannel log that states the cause).
The release asset address is https://github.com/flannel-io/flannel/releases/download/v0.28.9/kube-flannel.yml. The default Network in the manifest is 10.244.0.0/16, so you cannot use it as is on this VM. A node being Ready only means that a CNI configuration file has appeared, not that the daemon is alive. You can see the log of a container that is restarting with --previous.
Load br_netfilter to bring the CNI to life
Write the br_netfilter module to /etc/modules-load.d/k8s.conf so that it loads at boot, and load it now as well. Add net.bridge.bridge-nf-call-iptables = 1 to /etc/sysctl.d/k8s.conf and apply it. The Flannel DaemonSet must become ready, both CoreDNS Pods must become Available, and on the node, dig @<kube-dns 서비스 IP> kubernetes.default.svc.cluster.local (the placeholder is the kube-dns Service IP) must return the kubernetes Service IP.
CrashLoopBackOff increases the restart interval gradually, so you may have to wait a long time even after fixing the cause. A Pod managed by a DaemonSet is recreated right away even if you delete it. After loading the module, you must reread sysctl for the bridge entry to appear.
Control plane Pods that come back even when deleted
Write to /root/kubeadm/static-pods.json the fields manifests (a sorted array of the file names in /etc/kubernetes/manifests), owner_kind (the ownerReferences kind of the kube-scheduler Pod), etcd_data_dir (the data path the etcd manifest mounts as a hostPath), and service_cluster_ip_range (the value of the flag of the same name in the kube-apiserver manifest). Then delete the kube-scheduler Pod with kubectl delete, confirm that it comes back, and add deleted_uid, new_uid, container_id_before, and container_id_after (containerStatuses[0].containerID) to the same file.
A static Pod is started by the kubelet watching a directory, not by the API server. What you see in the API server is a mirror Pod created by the kubelet, so even if you delete it, the kubelet only registers the mirror again. Use the containerID to judge whether the container was restarted.
Pods do not get scheduled on a single-machine cluster
In the default namespace, create a Pod probe (image registry.k8s.io/pause:3.10.2, restartPolicy Never), confirm that it is Pending, and write to /root/kubeadm/pending.json the fields pod_uid, phase, message (the message of the PodScheduled condition), and observed_at (UTC, in the 2026-01-01T00:00:00Z format). Then remove the node-role.kubernetes.io/control-plane:NoSchedule taint from the node so that probe becomes Running. Do not delete and recreate the Pod.
kubeadm puts a taint on the control plane node so that ordinary Pods do not come to it. If there is only one node, that decision means 'nowhere to go.' The syntax for removing a taint is to append a minus sign after key:effect. The scheduler retries Pending Pods when the taint changes.
Prepare for one year from now and for a second node
Check with kubeadm certs check-expiration and openssl, and write to /root/kubeadm/certs.json the fields apiserver_not_after and ca_not_after (both UTC, in the 2026-01-01T00:00:00Z format), and apiserver_valid_days and ca_valid_days (the number of days from notBefore to notAfter of each certificate, an integer). Then create a new token with kubeadm token create --ttl 3h --description worker-join, calculate the CA public key hash yourself, and write to /root/kubeadm/join.txt a single line kubeadm join <API 엔드포인트> --token <토큰> --discovery-token-ca-cert-hash sha256:<해시> (the placeholders are the API endpoint, the token, and the hash).
The default lifetimes of the leaf certificates and the CA are set by certificateValidityPeriod and caCertificateValidityPeriod in the kubeadm configuration. The hash is the sha256 of the CA certificate's public key extracted as DER. The endpoint must match the server address written in the cluster-info ConfigMap in kube-public.
Report what you chose yourself
Write to /root/kubeadm/report.json the fields kubernetes_version (the server gitVersion), pod_cidr, service_cidr, dns_service_ip (the kube-dns Service IP), cgroup_driver (the driver the kubelet uses), notready_cause (the cause of the NotReady in step 3: one of cni · runtime · certificate), flannel_blocker (the name of the kernel module that blocked flannel in step 4), and join_token_id (the first 6 characters of the token in step 8).
Write them based on the records you left in earlier steps and the current cluster. The grader recalculates the same values from the cluster and from the records of earlier steps. You can check the kubelet's driver in the 'cgroup driver setting received from the CRI runtime' line of the kubelet log or in /var/lib/kubelet/config.yaml.