TT Lab
Get started
Learn Learning paths Courses

Building clusters with Kubespray and Terraform

failed=0 is only the start — checking a Kubespray cluster layer by layer

Continue in TT Lab

One-line summary

A cluster built with kubespray looks the same as a kubeadm cluster on the outside, but etcd, DNS, and certificates are placed differently in three spots, so verification must also look at those three separately.

Why this was needed

When the installation is done, people look at the single line kubectl get nodes and call it finished if it says Ready. But node Ready only means that the kubelet reports its status to the API server and that a CNI configuration file exists. Whether a Pod can pull an image, whether requests reach a Service IP, whether names resolve, and what will stop a year from now are not in that one line. Moreover, kubespray makes a few choices that differ from the kubeadm defaults. If you do not know those differences, you follow the kubeadm documentation's check procedure as is and skip the places that really matter.

How it works

Nodes and taints. kubeadm puts the node-role.kubernetes.io/control-plane:NoSchedule taint on the control plane node. If that node is also in kube_node in the inventory, kubespray removes this taint with the task "Remove taint for control plane node with node role". So a single-node kubespray cluster accepts ordinary Pods without any extra handling. The group placement in the inventory becomes the scheduling policy.

etcd is not a Pod. The default of etcd_deployment_type in group_vars/all/etcd.yml is host, so etcd comes up as the host's etcd.service from the binary kubespray downloaded. However much you look at kube-system, there is no etcd Pod, and you check its state with systemctl status etcd and /etc/etcd.env. The API server attaches to this etcd as an "external etcd".

Certificates come in two branches. What kubeadm creates (the API server, the controller and scheduler kubeconfigs, and so on) is in /etc/kubernetes/ssl in kubespray, with leaf certificates of 1 year and a CA of 10 years. On this VM, kubeadm certs check-expiration showed 364 days for leaf certificates and 9 years for the CA, and etcd was not in the list — because it is an external etcd. The etcd certificates are created with openssl by kubespray's make-ssl-etcd.sh, and their duration is certificates_duration (default 36500 days, about 100 years). If you look at which expires first, it becomes clear that the job within a year is renewal on the kubeadm side.

DNS has two tiers. kubespray turns on nodelocaldns by default. On each node a DaemonSet runs a caching DNS at a link-local address (default 169.254.25.10), and the kubelet's clusterDNS tells Pods that address. Only names the cache does not know pass on to CoreDNS. So to confirm that "DNS works", you must ask both the CoreDNS Service address and the nodelocaldns address, and the one Pods actually use is the latter. As the Kubernetes documentation's NodeLocal DNSCache explanation says, this arrangement is meant to reduce conntrack contention and CoreDNS load.

Records of the installation. kubespray leaves on the node the /etc/kubernetes/kubeadm-config.yaml it used when calling kubeadm, etcd's /etc/etcd.env, and the downloaded binary cache /tmp/releases (local_release_dir). Three months later, when someone asks "what values was this cluster built with?", these are the primary sources you look at after the inventory.

What it looks like in the field

On this course's measured cluster (1.35.8, calico), what came up in kube-system was calico-kube-controllers, calico-node, coredns, dns-autoscaler, kube-apiserver, kube-controller-manager, kube-proxy, kube-scheduler, and nodelocaldns. There was one CoreDNS, because the dns-autoscaler decides the replica count according to the number of nodes and cores. If nodes increase, CoreDNS increases too.

The most common misconception in the field is about certificates. If monitoring scrapes only the result of kubeadm certs check-expiration, the etcd certificates are never monitored. In this deployment the etcd side is 100 years so it is not a problem, but the story changes if someone has reduced certificates_duration or switched etcd to the kubeadm arrangement. On a checklist, "which tool created which certificate" must come first.

The last piece of evidence is the workload. If you send requests to a Deployment of two Pods and one Service and both Pod names come back, registry egress, scheduling, the Pod network, and kube-proxy are all confirmed at once. One request says more than several green lights.

What you will do in the next lab

After setting up the cluster, you read the node's roles and taints, and confirm from systemd why etcd is missing from the core Pod list. You ask nodelocaldns and CoreDNS each for a name to confirm the two tiers of DNS, and read the expiration dates of the kubeadm certificates and the etcd certificate with openssl. After checking the configuration files kubespray left, you finish by sending real requests to a small workload.