Building clusters with Kubespray and Terraform
Add-ons are group_vars switches — once on, prove they work
Goal
Before installation, you turn on metrics-server, local-path-provisioner, and Helm through group_vars, and after installation check that each one actually works.
You also look at the trap in which -e passes booleans as strings, and with which encapsulation the default network plugin, calico, is running.
Why it matters
In kubespray an add-on is a single variable. So turning it on is easy, and the fact that you turned it on and the fact that it works easily get mixed up. Even if metrics-server is up, if the APIService is not Available, HPA stops, and even if a storage class exists, if the PVC does not attach to the node disk, a stateful app does not come up. The network plugin is the hardest choice to change after installation, so you have to choose knowing what the default is and what network conditions that default requires.
Steps
- In
/root/ks/inventory/lab/group_vars/k8s_cluster/addons.yml, changemetrics_server_enabled,local_path_provisioner_enabled, andhelm_enabledall totrue. When resolved withansible-inventory --host node1, the three values must be boolean true, not strings. - In
/opt/ks/kubespray, runansible-playbook -i /root/ks/inventory/lab/inventory.ini playbooks/boilerplate.yml -e helm_enabled=trueand-e '{"helm_enabled": true}'separately. Write to/root/ks/addons/string-trap.jsonkv_failed_task(the name of the task at which the first command failed, without the role prefix),kv_type(the type helm_enabled became in the first command:"str"or"bool"), andjson_passed(whether the second command ended with failed=0, a boolean). - In
/opt/ks/kubespray, runansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.ymland leave the entire output in/root/ks/logs/cluster-1.log(about 8 minutes). node1 in the PLAY RECAP must havefailed=0. - Wait until
kubectl top nodeshows the CPU and memory of node1, and then write to/root/ks/addons/metrics.jsonapiservice_available(the status of the Available condition of the APIServicev1beta1.metrics.k8s.io),image(the container image of the metrics-server Deployment), andinsecure_tls(whether metrics-server is running with the--kubelet-insecure-tlsargument, a boolean). - In the default namespace, create a PVC
data(storageClassNamelocal-path, 64Mi, ReadWriteOnce) and a Podwriterthat attaches it at/data(imagebusybox:latest, commandsh -c 'echo kubespray > /data/hello.txt && sleep 3600'). The PVC must become Bound, andhello.txtmust exist with the contentkubesprayin that volume's directory under the node's local-path storage path. - Create a chart with
helm create /root/ks/addons/demoand install it withhelm install demo /root/ks/addons/demo -n demo --create-namespace --wait. The releasedemomust be deployed and the Pod must be Ready. Write to/root/ks/addons/helm.jsonhelm_version(helm version --template '{{.Version}}') andkubespray_helm_version(the default helm_version of this kubespray version, with a v in front). - Write to
/root/ks/addons/cni.jsonplugin(the kube_network_plugin the inventory resolved),calico_version(the image tag of the calico-node container of the calico-node DaemonSet),vxlan_modeandipip_mode(the spec values ofcalicoctl.sh get ippool default-pool -o json), andpool_cidr(the cidr of that IP pool).
Notes
- kubespray v2.32.0 is ready at
/opt/ks/kubesprayand the inventory at/root/ks/inventory/lab/inventory.ini(kube_version 1.35.8). You do the installation yourself in step 3. - The add-on images are pulled from registry.k8s.io and docker.io. This VM can go out over public 80/443.
- Common mistake: putting quotes on it in addons.yml, like
"true". Some variables are blocked by validate_inventory, and some are quietly resolved differently in conditional expressions. - Common mistake: reinstalling metrics-server after seeing kubectl top fail for the first few tens of seconds. Wait for the first collection cycle.
- Documentation: Kubespray — choosing a CNI (k8s-cluster.yml) · Kubespray — Calico · Kubernetes — Resource metrics pipeline
Choose add-ons before installation
In /root/ks/inventory/lab/group_vars/k8s_cluster/addons.yml, change metrics_server_enabled, local_path_provisioner_enabled, and helm_enabled all to true. When resolved with ansible-inventory --host node1, the three values must be boolean true, not strings.
In the sample's addons.yml, three lines are written as false. In YAML, an unquoted true is a boolean and "true" is a string. In the JSON output of ansible-inventory, the two are distinguished as true and "true".
-e passes a string
In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini playbooks/boilerplate.yml -e helm_enabled=true and -e '{"helm_enabled": true}' separately. Write to /root/ks/addons/string-trap.json kv_failed_task (the name of the task at which the first command failed, without the role prefix), kv_type (the type helm_enabled became in the first command: "str" or "bool"), and json_passed (whether the second command ended with failed=0, a boolean).
From ansible-core 2.19, conditional expressions must be booleans, and -e key=value always passes a string. kubespray checks the type of a few known booleans in validate_inventory. You can see the type with ansible -i ... node1 -m debug -a 'msg={{{{ helm_enabled | type_debug }}}}' -e helm_enabled=true.
Set up together with the add-ons
In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml and leave the entire output in /root/ks/logs/cluster-1.log (about 8 minutes). node1 in the PLAY RECAP must have failed=0.
Add-ons are installed in the last play of cluster.yml (Install Kubernetes apps). On an already built cluster, you can also rerun only that part with --tags apps or an add-on tag (metrics_server and so on) after changing group_vars, but here you turn them on and build from the start. Start it with systemd-run or tmux so that it keeps running even if the console is cut off.
Does metrics-server really produce figures
Wait until kubectl top node shows the CPU and memory of node1, and then write to /root/ks/addons/metrics.json apiservice_available (the status of the Available condition of the APIService v1beta1.metrics.k8s.io), image (the container image of the metrics-server Deployment), and insecure_tls (whether metrics-server is running with the --kubelet-insecure-tls argument, a boolean).
metrics-server attaches to the API server through the aggregation layer. If the APIService is not Available, kubectl top fails with 'Metrics API not available.' It takes several tens of seconds until the first figures appear. The kubespray default (metrics_server_kubelet_insecure_tls: true) means it does not verify the kubelet's serving certificate, so in production you should consider issuing the kubelet serving certificates properly.
The PVC attaches to the node disk
In the default namespace, create a PVC data (storageClassName local-path, 64Mi, ReadWriteOnce) and a Pod writer that attaches it at /data (image busybox:latest, command sh -c 'echo kubespray > /data/hello.txt && sleep 3600'). The PVC must become Bound, and hello.txt must exist with the content kubespray in that volume's directory under the node's local-path storage path.
local-path-provisioner creates a directory on the node disk and hands it out as a PV when a Pod that uses the PVC is scheduled (WaitForFirstConsumer). The PV's spec.hostPath.path or spec.local.path is that directory. busybox is the helper image that local-path uses, so kubespray has already pulled it.
One release with the Helm kubespray installed
Create a chart with helm create /root/ks/addons/demo and install it with helm install demo /root/ks/addons/demo -n demo --create-namespace --wait. The release demo must be deployed and the Pod must be Ready. Write to /root/ks/addons/helm.json helm_version (helm version --template '{{.Version}}') and kubespray_helm_version (the default helm_version of this kubespray version, with a v in front).
kubespray's Helm is a version pinned by checksum, downloaded from get.helm.sh and placed at /usr/local/bin/helm. The default version is the first key of helm_archive_checksums in roles/kubespray_defaults/vars/main/checksums.yml. The chart that helm create makes uses the nginx image from docker.io.
How is calico running
Write to /root/ks/addons/cni.json plugin (the kube_network_plugin the inventory resolved), calico_version (the image tag of the calico-node container of the calico-node DaemonSet), vxlan_mode and ipip_mode (the spec values of calicoctl.sh get ippool default-pool -o json), and pool_cidr (the cidr of that IP pool).
kubespray wraps calicoctl as /usr/local/bin/calicoctl.sh. The default IP pool is created from kube_pods_subnet, and the encapsulation is decided by the calico_vxlan_mode and calico_ipip_mode variables. VXLAN only needs L3 to pass between nodes, while IPIP needs IP protocol 4 to pass — the cloud firewall is the point that decides it.