Building clusters with Kubespray and Terraform
Can this server take a cluster before you install?
Goal
Before putting Kubernetes on a server used by a previous administrator, you check and fix swap, kernel modules, sysctl, and ports. And by reading the role code, you sort out which of these kubespray does by itself and which a person must clear away first.
Why it matters
On a freshly provisioned VM it looks as if kubespray does everything for you. In fact, turning off swap, loading br_netfilter, and turning on ip_forward are done by the preinstall and node roles. But servers in the field are not new. If a monitoring agent is holding 10250, kubespray pushes the installation through to the end as it is, and the kubelet cannot take the port, so the node does not come up. If a sysctl file left by a security audit is later in name order, everything is fine on the day of installation and Pod communication breaks at the first reboot. Both are the kind that shows up only after the whole long installation has run, so the value you look at once before installing is the cheapest.
Steps
- Before fixing anything, write to
/root/ks/preflight/before.jsonswap_total_kb(SwapTotal in /proc/meminfo, a number),ip_forward(the current kernel value, a number),br_netfilter(whether the module is loaded, a boolean), andbusy_ports(among the control plane ports 2379, 2380, 6443, 10250, 10257, and 10259, the ports that something is currently LISTENing on, an ascending array of numbers). - Turn off all swap that is on, and remove the swap line from
/etc/fstabor block it with a comment so that it is not turned on again even after a reboot. The swap file itself may be deleted or kept. - Load
overlayandbr_netfilternow, and write the two names one per line in/etc/modules-load.d/k8s.confso that they are loaded at boot too. - Write
net.ipv4.ip_forward = 1,net.bridge.bridge-nf-call-iptables = 1, andnet.bridge.bridge-nf-call-ip6tables = 1in/etc/sysctl.d/k8s.confand apply them. The current kernel values must be 1 for all three, and they must still be 1 for all three even if sysctl.d is reread in name order at reboot. - Find the port among the control plane ports that something is holding, stop the systemd unit that started it, and disable it so that it does not start again at boot. Write to
/root/ks/preflight/ports.jsonport(the port that was held, a number),unit(the name of the unit that started that process, including .service), andpid(the PID of that process before stopping it, a number). When you are done, none of 2379, 2380, 6443, 10250, 10257, or 10259 may be LISTENing. - In
/opt/ks/kubespray, gather this node's facts withansible -i /root/ks/inventory/lab/inventory.ini node1 -m setupand write to/root/ks/preflight/facts.jsonmemtotal_mb,processor_vcpus,distribution,distribution_version,kernel, anddefault_ipv4(the IPv4 address of the default route). Then find the control plane minimum memory in kubespray'sroles/kubernetes/preinstall/defaults/main.ymland write it together asminimal_master_memory_mb. - Read
roles/kubernetes/preinstallandroles/kubernetes/nodeof kubespray v2.32.0 and write to/root/ks/preflight/who-fixes.json, for each of the four problems,"kubespray"if kubespray fixes it by itself and"operator"if a person must clear it away first. The keys areswap,br_netfilter,ip_forward, andport_10250. And also writesysctl_file(the path of the file to which kubespray writes sysctl values, the default).
Notes
- kubespray v2.32.0 and the inventory
/root/ks/inventory/lab/inventory.ini(one node, local connection) are ready in the VM. This lab does not perform an installation. - Common mistake: doing only
swapoff -aand leaving fstab as is. It comes back on at the next boot. - Common mistake: only killing the process. A
Restart=alwaysunit starts it again right away. - Common mistake: finishing after looking only at your own sysctl file. Check the order in which
sysctl --systemreads. - Documentation: Kubernetes — Ports and Protocols · Kubernetes — Container Runtimes (prerequisites) · Kubespray — Port requirements
Write down the present before fixing
Before fixing anything, write to /root/ks/preflight/before.json swap_total_kb (SwapTotal in /proc/meminfo, a number), ip_forward (the current kernel value, a number), br_netfilter (whether the module is loaded, a boolean), and busy_ports (among the control plane ports 2379, 2380, 6443, 10250, 10257, and 10259, the ports that something is currently LISTENing on, an ascending array of numbers).
You can see LISTENing sockets with ss -ltnp. You can check whether a module is loaded with lsmod or /sys/module/<이름> (the placeholder is the module name). The grader compares against values measured separately when the VM was prepared, so if you write them after fixing, they will not match.
Turn swap off and keep it from coming back
Turn off all swap that is on, and remove the swap line from /etc/fstab or block it with a comment so that it is not turned on again even after a reboot. The swap file itself may be deleted or kept.
swapoff turns it off only for now. What turns swap on at boot is the line in fstab whose type is swap. Which file is swap is in swapon --show or /proc/swaps.
Kernel modules, both now and at boot
Load overlay and br_netfilter now, and write the two names one per line in /etc/modules-load.d/k8s.conf so that they are loaded at boot too.
containerd's default snapshotter uses overlay, and br_netfilter is what makes Pod traffic crossing a bridge go through iptables. modprobe is only for now, and modules-load.d is from the next boot.
With sysctl, the file read last wins
Write net.ipv4.ip_forward = 1, net.bridge.bridge-nf-call-iptables = 1, and net.bridge.bridge-nf-call-ip6tables = 1 in /etc/sysctl.d/k8s.conf and apply them. The current kernel values must be 1 for all three, and they must still be 1 for all three even if sysctl.d is reread in name order at reboot.
After running sysctl --system, read ip_forward again. The files in sysctl.d are read in name order, and for the same key the later one wins. The output of sysctl --system shows which files it read in which order. Whether to delete the previous administrator's file, rename it, or move your own file later is a judgment call.
What took the kubelet's place
Find the port among the control plane ports that something is holding, stop the systemd unit that started it, and disable it so that it does not start again at boot. Write to /root/ks/preflight/ports.json port (the port that was held, a number), unit (the name of the unit that started that process, including .service), and pid (the PID of that process before stopping it, a number). When you are done, none of 2379, 2380, 6443, 10250, 10257, or 10259 may be LISTENing.
ss -ltnp shows the PID, and the unit that started the PID can be seen in systemctl status <PID> or /proc/<PID>/cgroup. If you only kill the process, a unit with Restart=always starts it again right away. kubespray does not free up this port. The first kubeadm init is blocked by the Port-10250 error in preflight, but kubespray's retry puts that error in the ignore list and runs again, so the cause remains only in the first attempt higher up the log (control-plane/tasks/kubeadm-setup.yml).
This host as kubespray sees it
In /opt/ks/kubespray, gather this node's facts with ansible -i /root/ks/inventory/lab/inventory.ini node1 -m setup and write to /root/ks/preflight/facts.json memtotal_mb, processor_vcpus, distribution, distribution_version, kernel, and default_ipv4 (the IPv4 address of the default route). Then find the control plane minimum memory in kubespray's roles/kubernetes/preinstall/defaults/main.yml and write it together as minimal_master_memory_mb.
kubespray's pre-check (0040-verify-settings.yml) judges by these facts. If the memory is smaller than the minimum, the installation stops within a few minutes of starting. The output of the setup module is under ansible_facts, and the default route address is default_ipv4.address. If you do not give the ip variable separately, kubespray attaches the API server and etcd to this address.
What you leave to the tool and what a person looks at
Read roles/kubernetes/preinstall and roles/kubernetes/node of kubespray v2.32.0 and write to /root/ks/preflight/who-fixes.json, for each of the four problems, "kubespray" if kubespray fixes it by itself and "operator" if a person must clear it away first. The keys are swap, br_netfilter, ip_forward, and port_10250. And also write sysctl_file (the path of the file to which kubespray writes sysctl values, the default).
0010-swapoff.yml of preinstall removes the swap line from fstab, masks swap.target, and calls swapoff. You can find modules and sysctl in the node and preinstall roles, and the sysctl file path is sysctl_file_path in kubespray_defaults. There is no task anywhere that frees a port. As you experienced in step 4, also think about the fact that if there is a file that comes later in name order than the file kubespray writes, the value reverts at reboot.