Building clusters with Kubespray and Terraform
One step at a time, from the inventory — the rules of Kubespray upgrades
One-line summary
A kubespray upgrade is raising the version in the inventory by one step and running upgrade-cluster.yml, and since it does cordon → drain → upgrade → uncordon for each node, on a single-node cluster that period is the downtime itself.
Why this was needed
The Kubernetes project patches only the latest three minor versions. So you have to upgrade two or three times a year, and if you hold out without upgrading you have to skip several steps at once, which is not permitted. Kubernetes' Version Skew Policy says kube-apiserver must be raised only one minor version at a time, and the kubeadm upgrade documentation also says not to skip minor versions. kubespray adds one more rule of its own on top of this — raise the kubespray tag one step at a time too. Because the role defaults and the list of files to download change with each tag, if you jump straight from an old tag to the latest tag, nobody has tested what will break.
How it works
You raise the version in the inventory. The gist of the upgrade documentation is this. If you have written kube_version in the inventory, either fix that value before upgrade-cluster.yml or pass the new version with -e kube_version=..., and if you do neither, it "stays on the same version written in the inventory." Of the two methods, fixing the file is better. If you raise it only with -e, the cluster is on the new version but the inventory is on the old one, and the moment someone next runs cluster.yml with that inventory, it starts from a diverged state.
upgrade-cluster.yml keeps the order. This playbook is used only for an existing cluster, and it raises the control plane and etcd first and the workers later. For each node, the pre-upgrade role does cordon and drain, kubeadm upgrade replaces the static Pod manifests with the new version, and after the kubelet is raised, the post-upgrade role uncordons. The number of nodes raised at once is decided by serial (default 20%), and the documentation also tells you ways to stop and check at each node with upgrade_node_confirm and upgrade_node_pause_seconds. To upgrade nodes in batches, it says to first run facts.yml without a limit and then raise the control plane first with --limit "kube_control_plane:etcd".
Not everything follows upward. kubespray picks the versions of the etcd, CoreDNS, and pause images from a table per Kubernetes minor version (etcd_supported_versions, coredns_supported_versions, and so on). If two minor versions point to the same value in the table, that component stays as it is. This measurement was like that — both 1.35.8 and 1.36.4 had etcd 3.6.14, while for CoreDNS the versions the table points to differ, so it changed from 1.12.4 to 1.14.2. "We raised Kubernetes, so etcd must have gone up too" is just an assumption until you check.
The next step comes from the next tag. The highest version a tag accepts is the first key of the checksum list, and in v2.32.0 it is 1.36.4. If you have come up to this version, you cannot go higher with this kubespray. As in the "Multiple upgrades" section of the documentation, the next step is raising kubespray to the next tag (reinstalling Ansible with that tag's requirements.txt) and running upgrade-cluster.yml.
What it looks like in the field
While building this course, I measured 1.35.8 → 1.36.4 on the medium VM. The whole upgrade-cluster.yml took 368 seconds, of which the task in which kubeadm upgrades the first control plane took 94 seconds and the drain 16 seconds. Memory rose to a maximum of 1.56GiB, slightly higher than at installation (maximum 1.36GiB), but it was ample within 4GiB. The node UID stayed the same, and the drained web Pods came back with new UIDs after the uncordon. That is, on a single-node cluster, a drain is not "moving" but "stopping and starting again."
Two incidents often seen in the field. First, if a PodDisruptionBudget ties up all the replicas with minAvailable, the drain does not finish. The pre-upgrade role's default is drain_timeout: 360s, and after trying three times (drain_retries: 3) it fails, and because of upgrade_node_uncordon_after_drain_failure: true, it uncordons the node again and stops the playbook. The node is left on the old version. Second, if it gets past the drain but stops in a later step, the node is left cordoned. If you go home without running it again, the next day a node with scheduling blocked is discovered. That is why "is there any node that is unschedulable?" is at the top of the post-upgrade checklist.
What you will do in the next lab
You build with 1.35.8, bring up a workload, and take a picture before upgrading. You change the version in the inventory to 1.36.4 and run upgrade-cluster.yml, then take the same picture again and compare what happened to the node, the workload, etcd, and CoreDNS. You read cordon, drain, and uncordon in the log and judge whether the next step is possible with this kubespray.