Draining Nodes and Planning the Upgrade
Goal
You carry out by hand the procedure of safely emptying and restoring one node, and you turn the version skew rules and a canary rollout plan into verifiable artifacts rather than documents.
Why it matters
Most upgrade incidents happen not from not knowing the commands but from not knowing the order and the scope. Before an irreversible operation (kubeadm upgrade apply), an etcd snapshot must come first. Since control plane downgrade is not supported, the escape hatch when something goes wrong is only "restoring from a snapshot," not "going down." The scope question is handled by version skew. The kubelet may be lower than the apiserver but cannot be higher, and only kubectl allows one step above. If you do not know this asymmetry, you get stuck on questions like "I can use the latest kubectl, so why can't I do that with the kubelet?" Finally, the three stages of node work (cordon → drain → uncordon) and the PDB go together. Without a PDB, a drain can bring down a service entirely, and if the PDB is too tight, the drain never finishes.
Steps
- Create
/root/ops/upgrade/out/versions.json. The top-level keys areserver(the server version string inv1.x.yform) andnodes.nodeshas the three keyslab-node-0,lab-node-1, andlab-node-2, and the values must be each node's actual kubelet version string. - Cordon only
lab-node-1, and save the node list right afterward to/root/ops/upgrade/out/cordon.txt. The file must showlab-node-1andSchedulingDisabled, andlab-node-0must not be blocked. Then, in/root/ops/upgrade/out/cordon-note.txt, write one line saying that cordon does not touch Pods that are already running. - Create the namespace
ops-upgradeand in it create a Deployment that hasreplicas: 3, namedweb(Pod labelapp=web). Then create a PodDisruptionBudgetweb-pdbin the same namespace withspec.minAvailable: 2andspec.selector.matchLabels.app: web, and do not usemaxUnavailable. - Drain
lab-node-1and save its output to/root/ops/upgrade/out/drain.txt. After the drain finishes, leave which nodes the Pods of theops-upgradenamespace are on in/root/ops/upgrade/out/after-drain.txt, in three or more lines. This file must not containlab-node-1and must showlab-node-0orlab-node-2. - Uncordon
lab-node-1and save its output to/root/ops/upgrade/out/uncordon.txt. All three nodes must be schedulable, and the ready replicas of the Deploymentwebmust be back to 3. - After copying
/opt/lab/fixtures/k8s/skew-template.yamlto/root/ops/upgrade/skew.yaml, fill in the question marks underanswers. Leaveapiserver: "1.34"as is because it is the premise of the problem. The answer format is a minor version string like"1.31", and onlydowngrade_supportedis"yes"or"no". The keys to fill in aremin_kubelet,max_kubelet,max_kubectl,min_kubectl,max_controller_manager,next_upgrade_target, anddowngrade_supported. - Write
/root/ops/upgrade/runbook.mdin at least 500 bytes. Put each item on a different line and keep the following order. (1) etcd snapshot backup, (2)kubeadm upgrade plan, (3)kubeadm upgrade applyon the first control plane, (4)kubeadm upgrade nodeon the remaining nodes, and in the node-work sectiondrainmust appear beforeuncordon. Also write the constraint that you cannot skip minor versions as a sentence (for example: "one at a time"). - Attach the
upgrade.labhub.io/stagelabel to all three nodes.lab-node-0getscanary,lab-node-1getsbatch1, andlab-node-2getsbatch2. Then write the plan to/root/ops/upgrade/out/rollout.json.stagesis an array of length 3, and each element hasstage(canary/batch1/batch2) andnodes(an array of node names). The first stage iscanaryand must have exactly 1 node, and all three nodes must be assigned somewhere. Put averify_between_stageskey at the top level and write what to check between stages.
Reference
- The server version is in
kubectl version -o json, under.serverVersion.gitVersion, and a node's kubelet version is inkubectl get nodes -o json, under.items[].status.nodeInfo.kubeletVersion. Usingjq, withfrom_entries, you can easily build an object that uses the names as keys. - A drain refuses as is if there are DaemonSet Pods or Pods that use emptyDir. Add
--ignore-daemonsets,--delete-emptydir-data, and, if needed,--force. - You can see which node each Pod is on with
kubectl get pods -n ops-upgrade -o wide --no-headers, one per line. - If a label contains a dot and a slash, it needs escaping in jsonpath. The simplest way to check is
kubectl get nodes -L upgrade.labhub.io/stage. - Common mistake 1: writing
kubeadm upgrade applyandkubeadm upgrade nodetogether on one line in the runbook of step 7. The order is judged by the line number of each command's first appearance, so if they are on the same line it does not pass. The same goes fordrainanduncordon. - Common mistake 2: writing kubectl's upper bound as the same value as the apiserver in step 6. kubectl alone allows one step above.
- The lab Pod comes up fresh for each lab, so the cluster state from the earlier lab does not remain. Create the
ops-upgradenamespace and the Deployment yourself in step 3. This is exactly why operational procedures must be left in runbooks and manifests rather than in memory.
Collect the control plane and node versions
Create /root/ops/upgrade/out/versions.json. The top-level keys are server (the server version string in v1.x.y form) and nodes. nodes has the three keys lab-node-0, lab-node-1, and lab-node-2, and the values must be each node's actual kubelet version string.
An upgrade starts with knowing the current version exactly. The server version and each node's kubelet version are read from different places. The node information is under status.nodeInfo.
Block scheduling on just one node
Cordon only lab-node-1, and save the node list right afterward to /root/ops/upgrade/out/cordon.txt. The file must show lab-node-1 and SchedulingDisabled, and lab-node-0 must not be blocked. Then, in /root/ops/upgrade/out/cordon-note.txt, write one line saying that cordon does not touch Pods that are already running.
The principle is one at a time. Leave a note that cordon blocks only future scheduling and does not touch Pods that are already running.
Create a PDB to protect availability
Create the namespace ops-upgrade and in it create a Deployment that has replicas: 3, named web (Pod label app=web). Then create a PodDisruptionBudget web-pdb in the same namespace with spec.minAvailable: 2 and spec.selector.matchLabels.app: web, and do not use maxUnavailable.
You declare the minimum number of the 3 replicas that must stay alive. The minimum value and the maximum unavailable value cannot be used together.
Empty the node and check the Pods moved
Drain lab-node-1 and save its output to /root/ops/upgrade/out/drain.txt. After the drain finishes, leave which nodes the Pods of the ops-upgrade namespace are on in /root/ops/upgrade/out/after-drain.txt, in three or more lines. This file must not contain lab-node-1 and must show lab-node-0 or lab-node-2.
A drain evicts the existing Pods in addition to a cordon. It is refused without the options related to DaemonSets and emptyDir. After emptying, save which nodes the Pods are on.
Put the finished node back
Uncordon lab-node-1 and save its output to /root/ops/upgrade/out/uncordon.txt. All three nodes must be schedulable, and the ready replicas of the Deployment web must be back to 3.
If you forget uncordon, that node quietly sits idle. All three nodes must be schedulable, and the workload must also return to its original count.
Fill in the version skew table
After copying /opt/lab/fixtures/k8s/skew-template.yaml to /root/ops/upgrade/skew.yaml, fill in the question marks under answers. Leave apiserver: "1.34" as is because it is the premise of the problem. The answer format is a minor version string like "1.31", and only downgrade_supported is "yes" or "no". The keys to fill in are min_kubelet, max_kubelet, max_kubectl, min_kubectl, max_controller_manager, next_upgrade_target, and downgrade_supported.
The reference is always the apiserver. The kubelet is 3 minors below, the controllers are 1 minor below, and only kubectl is 1 minor in either direction. You raise a minor one at a time.
Write the upgrade runbook
Write /root/ops/upgrade/runbook.md in at least 500 bytes. Put each item on a different line and keep the following order. (1) etcd snapshot backup, (2) kubeadm upgrade plan, (3) kubeadm upgrade apply on the first control plane, (4) kubeadm upgrade node on the remaining nodes, and in the node-work section drain must appear before uncordon. Also write the constraint that you cannot skip minor versions as a sentence (for example: "one at a time").
A backup comes before an irreversible operation. Also write down that the kubeadm subcommands used by the first control plane and by the remaining nodes are different.
Make a canary rollout plan and attach labels
Attach the upgrade.labhub.io/stage label to all three nodes. lab-node-0 gets canary, lab-node-1 gets batch1, and lab-node-2 gets batch2. Then write the plan to /root/ops/upgrade/out/rollout.json. stages is an array of length 3, and each element has stage (canary/batch1/batch2) and nodes (an array of node names). The first stage is canary and must have exactly 1 node, and all three nodes must be assigned somewhere. Put a verify_between_stages key at the top level and write what to check between stages.
Do not leave the plan only as a document; engrave it into the cluster with node labels. The first stage must be a single machine, and what you check between stages is the core of a canary.