Version Skew and the Order of a Zero-Downtime Upgrade
Summary
An upgrade is not "installing a new version" but "moving the components forward one step at a time while keeping their version differences within the allowed range."
Why this matters
Kubernetes is not one program but a distributed system in which the apiserver, controller-manager, scheduler, kubelet, kube-proxy, and kubectl each run as a separate process. It is impossible to replace them all at once. With seven nodes there are seven kubelets, and if one of them fails to reboot, only that node is left on the old version. That is why Kubernetes officially defines "the range within which versions may differ." This is the version skew policy.
This problem also appears in the author's homelab expansion record. The two newly attached mini PCs still had a v1.32 kubelet left over from an old cluster, and if they had joined in that state, the control plane and the nodes would have run with mismatched versions. In the end it was tidied up with a script that starts over from teardown. A version is not a value that only has to be roughly similar, but a contract whose supported range is fixed in documentation.
How it works
The reference point is always kube-apiserver. The allowed range of the others is set around the apiserver.
| Component | Allowed range relative to the apiserver |
|---|---|
| kubelet | May be up to 3 minors lower. Must never be higher |
| kube-proxy | Same rule as the kubelet |
| controller-manager / scheduler | May be up to 1 minor lower. Must not be higher |
| kubectl | Up to 1 minor in either direction (±1) |
Say the apiserver is 1.34. Applying this table as is, the kubelet can be 1.31 through 1.34, the controller-manager at most 1.34, and kubectl 1.33 through 1.35. The point that kubectl alone can be higher than the apiserver is a spot people often get confused about.
Two more rules come with this.
- One minor at a time. You cannot go directly from 1.32 to 1.34. You must pass through 1.33. If you skip, the conversion path for the stored API objects is not guaranteed.
- Control plane downgrade is not supported. During an upgrade, the data in etcd is written in the new schema, so the way back is not "going down to the previous version" but restoring from the etcd snapshot taken just before the upgrade. This one sentence connects the previous module to this one.
The actual order for a kubeadm cluster is this.
etcd 스냅샷 백업
→ kubeadm upgrade plan (무엇이 가능한지 확인)
→ 첫 컨트롤 플레인에서 kubeadm upgrade apply v1.x.y
→ 나머지 컨트롤 플레인에서 kubeadm upgrade node
→ 노드마다: drain → kubelet/kubeadm 패키지 교체 → uncordon
The difference between drain and cordon is decided here too. cordon only blocks scheduling from now on and leaves the Pods that are already running as they are. drain additionally evicts the existing Pods. So the standard order for node work is to cut off inflow with cordon, empty it with drain, do the work, and put it back with uncordon. If you forget uncordon, that node quietly sits idle.
The mechanism that keeps a service from dying while Pods are evicted is the PodDisruptionBudget. If you set minAvailable: 2, an evict request is approved only as long as it does not cross that line. However, you can use only one of minAvailable and maxUnavailable.
What it looks like in the field
First, a bulk upgrade with no canary. If you upgrade all the nodes at once, there is no control group to go back to when something goes wrong. The basic approach is to upgrade one node first, confirm the workloads are healthy, and then split the rest into batches. If you attach stage labels to the nodes, the plan becomes cluster state rather than a document.
Second, a PDB that blocks the drain forever. If you have only 1 replica and set minAvailable: 1, that Pod can never be evicted. If a drain has been stuck for several minutes, look at the PDB first.
Third, "the status is Ready" and "it actually works" are different propositions. The author recorded the time when they installed KubeVirt and all the components were AllComponentsReady but the VMs would not come up. The same applies after an upgrade. That a node is Ready and that the workloads are healthy must be checked separately.
What to do in the next lab
You cordon a real node, move a workload with a PDB off it using drain, and put it back with uncordon. Using apiserver 1.34 as the reference, you fill in the blanks of the skew table yourself, write the upgrade runbook from backup to uncordon, and attach stage labels to the nodes to build a canary rollout plan as JSON.