Building clusters with Kubespray and Terraform
Why you choose a network plugin, and what to check after enabling add-ons
One-line summary
kubespray's add-ons and network plugin are switches in group_vars, turning a switch on and that thing actually working must be checked separately, and only the network plugin is hard to change later, so choose it before installation with a reason.
Why this was needed
kubeadm picks neither a CNI, nor metrics-server, nor storage. A person has to find the manifests and apply them with matching versions, and a person also has to check whether those versions match the cluster version. kubespray bundled these into variables. The single line kube_network_plugin: calico creates calico's CRDs, DaemonSet, and IP pool, and the single line metrics_server_enabled: true creates metrics-server and the APIService. The versions and image addresses are pinned by the kubespray release, so the same inventory installs the same add-ons whenever you run it. In exchange, only the fact that you turned it on is recorded, and checking that it works is still a person's job.
How it works
The network plugin. It is chosen by kube_network_plugin in group_vars/k8s_cluster/k8s-cluster.yml. The choices the sample's comments mention are cilium, calico, kube-ovn, flannel, and, for when you install it yourself, cni (only unpacks the plugin binaries and leaves the configuration blank) and none. The default is calico, and the calico defaults in v2.32.0 are calico_vxlan_mode: Always and calico_ipip_mode: Never. That is, it wraps Pod traffic between nodes in VXLAN (UDP 4789).
There are usually three reasons for choosing. First, NetworkPolicy — flannel gives only a Pod network and does not enforce policy, so if you need policy, it is calico or cilium. Second, network conditions — IPIP needs IP protocol 4 to pass between nodes, VXLAN needs UDP 4789, and BGP mode needs TCP 179. This is why kubespray's port documentation has this table separately for each CNI. Third, operational burden — cilium replaces kube-proxy with eBPF and even provides observability (Hubble), but the kernel version and the configuration items increase. Either way, changing it after installation is effectively a reinstall. This is because the Pod IP range and the node routing are tied to the plugin.
Add-ons. They are gathered in group_vars/k8s_cluster/addons.yml and in the sample all of them are false. The three turned on in this module are these.
metrics_server_enabled kubectl top · HPA 의 자원 지표. aggregation layer 로 API 서버에 붙는다
local_path_provisioner_enabled 노드 디스크(/opt/local-path-provisioner/)를 PV 로 주는 동적 프로비저너
helm_enabled helm 바이너리를 체크섬으로 고정한 판으로 /usr/local/bin 에 둔다
The add-ons are installed in the last play of cluster.yml, "Install Kubernetes apps", and each role has a tag (metrics_server and so on), so you can also rerun only that part on an already built cluster.
The boolean trap. From ansible-core 2.19 (Ansible 12), conditional expressions must be booleans. But -e helm_enabled=true passes the string "true". kubespray blocks a few well-known switches with "Stop if known booleans are set as strings" in validate_inventory, but when measured, this check looks at only four: download_run_once, download_always_pull, helm_enabled, and openstack_lbaas_enabled. If you give the other switches as strings, they do not get caught by that check and become a separate problem where the conditional expression is used. The answer is, as the release notes recommend, to pass it as JSON like -e '{"helm_enabled": true}', or to write it in the inventory as an unquoted YAML boolean.
What it looks like in the field
It is common to turn on metrics-server and then get an inquiry saying "HPA doesn't scale up." The order of checking is the APIService's Available condition → the metrics-server log → the connection to the kubelet. The kubespray default metrics_server_kubelet_insecure_tls: true means that metrics-server does not verify the kubelet's serving certificate. It is convenient in a single-node lab, but in production you should consider issuing the kubelet serving certificates properly (kubelet_rotate_server_certificates) and turning this value off. A trade like this is usually hidden in easy-to-enable defaults.
local-path-provisioner creates a directory on that node's disk when a Pod that uses a PVC is scheduled. It is convenient, but the data is tied to the node, so if the node disappears the data disappears too. It is good for labs and caches and does not suit data that needs replication.
What you will do in the next lab
You turn on three add-ons with addons.yml, and check how -e key=value and the JSON form are resolved differently by the boilerplate check. After installing, you check in turn kubectl top and the APIService, PVCs attaching to the node disk, and a release made with the Helm that kubespray installed, and read the encapsulation method from calico's IP pool.