TT Lab
Get started
Learn Learning paths Courses

CKA — Kubernetes Administrator

Three nodes — join, drain, kubelet recovery

Continue in TT Lab

Goal

You ssh into three nodes to assemble a kubeadm cluster, drain one worker for maintenance, and then revive the other worker, which does not come back after a reboot. You practice, on three real VMs, the "hands that work across several nodes" that the CKA hands-on exam requires.

Why it matters

On a single-machine cluster, you cannot see what a node is. The moment there are several nodes, three things change. First, the node's address becomes a problem. In this lab's three VMs, the first NIC's address is 10.0.2.2 on all of them, and the real cluster network is the second NIC (eth1). The kubelet's --node-ip, the API server's advertise address, and the CNI's tunnel NIC must all point to eth1 — this is actually the place people get wrong most often when setting up Kubernetes on a multi-NIC server. Second, maintenance comes up. Draining Pods before shutting a node down, and putting it back with uncordon, are everyday operations. Third, the kubelet is not a Pod but a systemd service. When a node is NotReady, you cannot reach the cause with kubectl, and you have to go into that node and look at the service.

Steps

  1. From controlplane, use ssh node01 and ssh node02 to get into each node and check hostname and the IPv4 address of eth1. Write the three nodes to /root/cluster/nodes.txt, one per line as 이름 주소 (name and address), in the order controlplane, node01, node02.
  2. On controlplane, run kubeadm init with --kubernetes-version v1.36.4 --apiserver-advertise-address <controlplane 의 eth1 주소> --pod-network-cidr 172.20.0.0/16 --service-cidr 172.21.0.0/16 (where the placeholder is the eth1 address of the controlplane), and copy /etc/kubernetes/admin.conf to /root/.kube/config. The controlplane node's InternalIP must be the eth1 address.
  3. Take the pre-downloaded /root/cluster/kube-flannel.yml (Flannel v0.28.9) and, in the net-conf.json section, set Network to 172.20.0.0/16; in the args of the kube-flannel container, add --iface=eth1, and apply it. The controlplane must become Ready and CoreDNS must be Available.
  4. On controlplane, generate the join command (kubeadm token create --print-join-command) and run it on node01 and node02. All three nodes must be Ready, and the InternalIP of node01 and node02 must be each one's own eth1 address.
  5. In the default namespace, create the Deployment web with the image registry.k8s.io/e2e-test-images/agnhost:2.53, the command /agnhost netexec --http-port=8080, and 2 replicas. Both Pods must be Ready on the workers, and from the controlplane, :8080/hostname on each Pod IP must return that Pod's name.
  6. Drain node01 with kubectl drain (leave the DaemonSet Pods) and save the entire output to /root/cluster/drain.log. node01 must have no Pods that are not from a DaemonSet, and both Pods of web must be Ready.
  7. Reboot node02 (ssh node02 reboot). After it comes back, find on node02 the cause of node02 staying NotReady and fix it, and make it come back by itself on the next reboot too.
  8. Make node01 schedulable again. All three nodes must be Ready and schedulable, the Flannel DaemonSet must be ready on all three nodes, and both Pods of web must be Ready.

Reference

Get into the three nodes and check the second NIC

From controlplane, use ssh node01 and ssh node02 to get into each node and check hostname and the IPv4 address of eth1. Write the three nodes to /root/cluster/nodes.txt, one per line as 이름 주소 (name and address), in the order controlplane, node01, node02.

ip -4 addr shows two NICs. Check that the first NIC's (enp1s0) address is the same on all three nodes — you cannot tell the nodes apart by that address. You can also just send the command, as in ssh node01 'ip -4 -o addr show eth1'.

Bring up the control plane with the eth1 address

On controlplane, run kubeadm init with --kubernetes-version v1.36.4 --apiserver-advertise-address <controlplane 의 eth1 주소> --pod-network-cidr 172.20.0.0/16 --service-cidr 172.21.0.0/16 (where the placeholder is the eth1 address of the controlplane), and copy /etc/kubernetes/admin.conf to /root/.kube/config. The controlplane node's InternalIP must be the eth1 address.

If you do not give an advertise address, kubeadm picks the NIC of the default route (10.0.2.2). Then the workers go to that address to find the API server and connect to themselves. The address on the kubelet side is set in /etc/default/kubelet, by --node-ip — it is already written there, so read it.

Tell Flannel which NIC to build the tunnel on

Take the pre-downloaded /root/cluster/kube-flannel.yml (Flannel v0.28.9) and, in the net-conf.json section, set Network to 172.20.0.0/16; in the args of the kube-flannel container, add --iface=eth1, and apply it. The controlplane must become Ready and CoreDNS must be Available.

Flannel builds VXLAN on the NIC of the default route. If that NIC's address is the same on all three nodes, every tunnel destination is itself — the symptom is that the nodes are Ready but cannot reach Pods on other nodes. Add the args as one more line under the --kube-subnet-mgr line, with the same indentation.

Join the two workers

On controlplane, generate the join command (kubeadm token create --print-join-command) and run it on node01 and node02. All three nodes must be Ready, and the InternalIP of node01 and node02 must be each one's own eth1 address.

You run the join command as root on the worker — ssh node01 '<조인 명령>' (where the placeholder is the join command). Even after the join finishes, Ready is delayed a little until the Flannel Pod comes up on that node. Read the WARNING lines in the join output too. They will be useful later.

Check that you can really reach Pods on other nodes

In the default namespace, create the Deployment web with the image registry.k8s.io/e2e-test-images/agnhost:2.53, the command /agnhost netexec --http-port=8080, and 2 replicas. Both Pods must be Ready on the workers, and from the controlplane, :8080/hostname on each Pod IP must return that Pod's name.

One line of kubectl create deployment web --image=… --replicas=2 -- /agnhost netexec --http-port=8080 is enough. The controlplane has the control-plane taint, so Pods go only to the workers. If curl hangs, suspect Flannel's NIC first.

Drain node01 for maintenance

Drain node01 with kubectl drain (leave the DaemonSet Pods) and save the entire output to /root/cluster/drain.log. node01 must have no Pods that are not from a DaemonSet, and both Pods of web must be Ready.

A drain is a cordon (scheduling block) plus eviction. DaemonSet Pods such as Flannel and kube-proxy come right back on the same node even if evicted, so you skip them with --ignore-daemonsets. The evicted web Pod moves to node02. You can keep the output on both the screen and a file with 2>&1 | tee.

The rebooted node02 does not come back

Reboot node02 (ssh node02 reboot). After it comes back, find on node02 the cause of node02 staying NotReady and fix it, and make it come back by itself on the next reboot too.

If a node is NotReady, look at that node's kubelet first — systemctl status kubelet. Whether it is running now (active) and whether it starts at boot (enabled) are different questions. Do you remember the WARNING that appeared when you joined?

Finish the maintenance and put it back

Make node01 schedulable again. All three nodes must be Ready and schedulable, the Flannel DaemonSet must be ready on all three nodes, and both Pods of web must be Ready.

The command that undoes a drain is uncordon. Even after you undo it, Pods that have already moved to node02 do not return on their own — the scheduler makes a decision only when it places a new Pod.