TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

What changes when there is more than one node

Continue in TT Lab

Summary

Once there are several nodes, the problem in Kubernetes changes from "the Pod won't come up" to "what is wrong on which node." The answer usually is not in kubectl but inside that node.

Why this matters

When you practice on a single-machine cluster, the node looks like background. With only one node name, there is nothing to think about for scheduling, networking, or maintenance. But a real cluster has several nodes, and the CKA hands-on exam also has you work by going into several nodes over ssh. Joining a worker, draining a node, and reviving a NotReady node all happen inside the node.

How it works

A node's address. The kubelet registers its own address as the node object's InternalIP, and the API server, other nodes, and the CNI find this node by that address. If you do not give an address separately, the kubelet picks the address of the NIC that the default route goes out of. On a server with one NIC, that is correct, but on a server whose management network and cluster network are split, the wrong address gets registered. That is why you have to answer the same question separately for each component, as with --node-ip (kubelet), --apiserver-advertise-address (kubeadm), and --iface (Flannel). In this lab's VMs, the first NIC is 10.0.2.2 on all of them, so if you miss even one place, the symptom shows up immediately.

Joining. kubeadm join introduces itself to the API server with a bootstrap token, confirms that API server is genuine with --discovery-token-ca-cert-hash, then receives the certificate the kubelet will use and writes /etc/kubernetes/kubelet.conf. When the join is done, the kubelet creates the node object, and the node becomes Ready only when the CNI DaemonSet brings up a Pod on that node.

Maintenance. kubectl cordon only marks the node as unschedulable. kubectl drain additionally evicts the Pods — unlike deletion, eviction respects the PodDisruptionBudget. DaemonSet Pods come right back on the same node even if you evict them, so you skip them with --ignore-daemonsets. When maintenance is finished you put the node back with uncordon, but Pods that have already moved do not come back. This is because the scheduler makes its decision only when it places a new Pod.

The kubelet is a service. The control plane components run as static Pods, but the kubelet itself, which brings up those static Pods, is managed by systemd. So if the kubelet dies, there is nothing to report that node's status, so the node becomes NotReady, and the cause remains only in systemctl status kubelet and journalctl -u kubelet. Here, active (running now) and enabled (started at boot) are different questions. When joining, kubeadm starts the kubelet directly but does not enable it, leaving only a warning.

What this lab's environment has in common with and differs from the real thing

Let's start with what is the same. The three VMs are real servers, each with its own kernel, systemd, and containerd, and traffic between nodes rides a real network. Flannel's VXLAN goes out over UDP 8472, the API server is called on 6443, and the kubelet on 10250. So if one setting is wrong, the symptom appears exactly as on a real cluster.

There are two differences. First, eth1 is not a real NIC but a virtual L2 network connecting only the VMs of this session. Seen from outside, only UDP 4789 passes between the VMs, and it does not reach other students' VMs. Second, the first NICs of all three VMs receive the same address, 10.0.2.2. This is rare on real servers, but thanks to it, "what gets chosen if you do not specify an address" shows up clearly rather than blurrily. A configuration that relies on defaults will certainly fail in this environment.

What it looks like in the field

On-prem servers commonly have separate NICs for the management, storage, and service networks. If you bring up kubeadm with defaults on them, the nodes register with their management network addresses, and Pod-to-Pod traffic rides the storage network switch or does not get through at all. All the nodes look Ready, so you find out only much later. It is also common for a node that rebooted after a kernel patch not to come back. Someone started the service by hand but did not enable it, and that fact shows no symptom until the reboot. That is why "check that it returns to Ready after the reboot" is always the last line of a maintenance runbook. Maintenance ends not at the moment you shut a node down but at the moment you confirm it has come back, and a reboot you do not verify is the same as scheduling the next outage.

What to do in the next lab

You ssh into the three nodes and check the NICs, bring up the control plane with the eth1 address, and tell Flannel which NIC to use for the tunnel. After you join the two workers and confirm you can actually reach Pods on other nodes, you drain node01 for maintenance and revive node02, which does not come back after a reboot, from inside that node.