Building clusters with Kubespray and Terraform
From 1.35.8 to 1.36.4 — same node, new version
Goal
With kubespray v2.32.0, you build a Kubernetes 1.35.8 cluster and bring up a workload, then change the version in the inventory to 1.36.4 and
raise it one step with upgrade-cluster.yml. You leave pictures before and after the upgrade and use them as evidence to say what changed and what stayed the same.
Why it matters
Kubernetes releases a minor version roughly every four months, and once the support period passes, security patches stop. So an upgrade is not a one-time job but a regular task. kubespray handles upgrades with the same declaration as installation — you change the version in the inventory and run the playbook. There is an order, though. You do not skip minors, you do not skip kubespray tags, and for each node you do cordon → drain → upgrade → uncordon. In this lab you also see for yourself that with one node, workloads stop during the drain. You wait about 13 minutes in total for the installation and the upgrade. If you are short on time, extend the session.
Steps
- In
/opt/ks/kubespray, runansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.ymland leave the entire output in/root/ks/logs/cluster-1.log(about 7 minutes). The inventory's kube_version is 1.35.8. When it finishes, the API server version must be v1.35.8. - In the default namespace, create a Deployment
web(imageregistry.k8s.io/e2e-test-images/agnhost:2.59, argumentsnetexec --http-port=8080, 2 replicas) and make both Pods Ready. Then write to/root/ks/upgrade/before.jsonserver_version,kubelet_version,etcd_version(the version ofetcd --version),coredns_image(the image of the coredns Deployment),node_uid, andweb_pod_uids(a sorted array of the web Pod UIDs). - Change
kube_versionin/root/ks/inventory/lab/group_vars/k8s_cluster/k8s-cluster.ymlto 1.36.4. Do not pass it with-e; edit the file.ansible-inventory --host node1must return 1.36.4. - In
/opt/ks/kubespray, runansible-playbook -i /root/ks/inventory/lab/inventory.ini upgrade-cluster.ymland leave the entire output in/root/ks/logs/upgrade.log(about 6 minutes). The PLAY RECAP must havefailed=0, the API server and the kubelet must both be v1.36.4, and the kubernetesVersion in/etc/kubernetes/kubeadm-config.yamlmust also be v1.36.4. - Write
/root/ks/upgrade/after.jsonwith the same fields as in step 2. And write to/root/ks/upgrade/diff.json, as booleans,same_node(whether the node UID stayed the same),web_pods_recreated(whether all the web Pod UIDs changed),etcd_changed(whether the etcd version changed), andcoredns_changed(whether the coredns image changed). - In
/root/ks/logs/upgrade.log, find the tasks of the upgrade/pre-upgrade and post-upgrade roles and write to/root/ks/upgrade/drain.jsoncordon_changed,drain_changed, anduncordon_changed(whether the "Cordon node", "Drain node", and "Uncordon node" tasks, respectively, reported changed, booleans),drain_sec(the seconds the "Drain node" task took, the TASKS RECAP value rounded to an integer, 0 if it is not in the list), andnode_schedulable_now(whether the node's spec.unschedulable is not true now, a boolean). - Write to
/root/ks/upgrade/next.jsonrunning(the current API server version),max_supported(the highest Kubernetes version this kubespray version can install, the first key of the checksum list), andneeds_new_kubespray(whether you must first raise kubespray to the next tag to go up to the next minor, a boolean).
Notes
- kubespray v2.32.0 is ready at
/opt/ks/kubesprayand the inventory at/root/ks/inventory/lab/inventory.ini(kube_version 1.35.8). - Common mistake: raising only with
-e kube_version=1.36.4and leaving the inventory as it is. The next person ends up running cluster.yml with an inventory that has the old version written in it. - Common mistake: not noticing that the node was left cordoned after an upgrade failed midway. Check SchedulingDisabled in
kubectl get node. - Documentation: Kubespray — Upgrading Kubernetes · Kubernetes — Version Skew Policy · Kubernetes — Upgrading kubeadm clusters
Build one step below (1.35.8)
In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini cluster.yml and leave the entire output in /root/ks/logs/cluster-1.log (about 7 minutes). The inventory's kube_version is 1.35.8. When it finishes, the API server version must be v1.35.8.
To practice upgrading, a version one step below must exist first. The default of this kubespray version is 1.36.4, so if you do not write 1.35.8 in the inventory, 1.36.4 is installed from the start.
The picture before upgrading
In the default namespace, create a Deployment web (image registry.k8s.io/e2e-test-images/agnhost:2.59, arguments netexec --http-port=8080, 2 replicas) and make both Pods Ready. Then write to /root/ks/upgrade/before.json server_version, kubelet_version, etcd_version (the version of etcd --version), coredns_image (the image of the coredns Deployment), node_uid, and web_pod_uids (a sorted array of the web Pod UIDs).
To say what changed and what stayed the same after the upgrade, you must first leave the values from before upgrading. etcd is a host service, not a Pod, so you ask its version from the binary, not from kubectl. The grader compares by version number whether this record was written before the upgrade.
Raise the version in the inventory
Change kube_version in /root/ks/inventory/lab/group_vars/k8s_cluster/k8s-cluster.yml to 1.36.4. Do not pass it with -e; edit the file. ansible-inventory --host node1 must return 1.36.4.
The kubespray upgrade documentation says that if you have written the version in the inventory, you should fix that value before running upgrade-cluster.yml. If you raise only with -e, the old version remains in the inventory, and when someone next runs cluster.yml with that inventory, it starts with the cluster and the inventory diverged.
upgrade-cluster.yml
In /opt/ks/kubespray, run ansible-playbook -i /root/ks/inventory/lab/inventory.ini upgrade-cluster.yml and leave the entire output in /root/ks/logs/upgrade.log (about 6 minutes). The PLAY RECAP must have failed=0, the API server and the kubelet must both be v1.36.4, and the kubernetesVersion in /etc/kubernetes/kubeadm-config.yaml must also be v1.36.4.
upgrade-cluster.yml is a playbook used only for an existing cluster, and it does cordon → drain → upgrade → uncordon for each node. If there is only one node, during the drain there is nowhere for the Pods to go, so they become Pending for a while — this means that upgrading a single-node cluster is itself downtime. While you wait, follow the PLAY headers in the log.
The picture after upgrading
Write /root/ks/upgrade/after.json with the same fields as in step 2. And write to /root/ks/upgrade/diff.json, as booleans, same_node (whether the node UID stayed the same), web_pods_recreated (whether all the web Pod UIDs changed), etcd_changed (whether the etcd version changed), and coredns_changed (whether the coredns image changed).
Raising Kubernetes one step does not mean every component follows upward. kubespray picks the etcd and CoreDNS versions from a table per Kubernetes minor version (etcd_supported_versions, coredns_supported_versions). If two versions point to the same value in the table, it stays as it is.
Reading the drain from the log
In /root/ks/logs/upgrade.log, find the tasks of the upgrade/pre-upgrade and post-upgrade roles and write to /root/ks/upgrade/drain.json cordon_changed, drain_changed, and uncordon_changed (whether the "Cordon node", "Drain node", and "Uncordon node" tasks, respectively, reported changed, booleans), drain_sec (the seconds the "Drain node" task took, the TASKS RECAP value rounded to an integer, 0 if it is not in the list), and node_schedulable_now (whether the node's spec.unschedulable is not true now, a boolean).
The TASKS RECAP list of profile_tasks shows only the tasks that took a long time. If the drain is not in the list, it means it finished quickly. If an upgrade fails midway, the node may be left cordoned, so you need the habit of checking unschedulable after it ends.
Where does the next step come from
Write to /root/ks/upgrade/next.json running (the current API server version), max_supported (the highest Kubernetes version this kubespray version can install, the first key of the checksum list), and needs_new_kubespray (whether you must first raise kubespray to the next tag to go up to the next minor, a boolean).
kubespray pins the versions it accepts for each tag by checksum. If the current version is at the top of the list, you cannot go higher with this kubespray, and as in the 'Multiple upgrades' section of the documentation, you must raise the kubespray tag one step at a time and then run upgrade-cluster.yml. Skipping tags is not supported.