The kernel moved up one notch, and only that node lost its driver
Goal
Taking node labels as the place of facts, you build a tool that makes the tag of a precompiled driver image and a checker that finds mismatched nodes, experience what is needed first when nodes grow, and, passing the spot a PodDisruptionBudget blocks, walk the upgrade procedure to the end, from drain to uncordon.
Why it matters
A GPU driver is not an application but a kernel module. Even when the GPU Operator brings up the driver as a container, what that container does is load the module into the host kernel, so when the kernel changes, the driver must also be rebuilt to match that kernel. That the tag of a precompiled driver image is <드라이버브랜치>-<커널판>-<OS태그> (the placeholders are the driver branch, the kernel version, and the OS tag) bakes this fact into the name. Two things follow from this. First, adding a node or raising the kernel by one is itself image work — if you pass without knowing, only on that node does the driver Pod die for lack of an image. Second, to raise the driver you must first bring down that node's GPU workloads, and since that goes through the eviction API, a PodDisruptionBudget blocks the upgrade. A good share of the reports "the upgrade stopped on one node" are budget problems, not driver problems, and if you do not know the cause, you end up reading driver logs for hours.
Steps
- Work in
/root/gpudrv(export KUBECONFIG=/root/.kube/config). First create the namespacegpu-drv(the Pods in later steps come up here). Attach labels to the three nodes. lab-node-0:nvidia.com/gpu.present=true,nvidia.com/cuda.driver.major=550,nvidia.com/cuda.driver.minor=90,nvidia.com/cuda.driver.rev=07,feature.node.kubernetes.io/kernel-version.full=5.15.0-119-generic,feature.node.kubernetes.io/system-os_release.ID=ubuntu, andfeature.node.kubernetes.io/system-os_release.VERSION_ID=22.04. lab-node-1: with the same keys, driver535/183/06, kernel5.15.0-107-generic,ubuntu22.04. lab-node-2: driver550/90/07, kernel6.8.0-45-generic,ubuntu24.04. All three nodes havenvidia.com/gpu.present=true. - Create
/root/gpudrv/drv-tag.sh <노드이름>(the placeholder is the node name). It reads that node's labels and prints the tag of a precompiled driver image as one line. The format is the official documentation's<드라이버브랜치>-<커널판>-<OS태그>(the placeholders are the driver branch, the kernel version, and the OS tag), where the branch isnvidia.com/cuda.driver.major, the kernel version isfeature.node.kubernetes.io/kernel-version.full, and the OS tag issystem-os_release.IDandsystem-os_release.VERSION_IDwritten together (for example,ubuntuand22.04giveubuntu22.04). Run it on the three nodes in turn and write three lines in/root/gpudrv/out/tags.txtin the form<노드> <태그>(the placeholders are the node and the tag). Do not write node names or tags inside the script — the grader calls it directly for each node. - Create
/root/gpudrv/support-matrix.csv. The first line is the headerdriver_branch,kernel,os_tag, followed by three lines of combinations actually built in the company registry —550,5.15.0-119-generic,ubuntu22.04,535,5.15.0-107-generic,ubuntu22.04, and550,6.8.0-45-generic,ubuntu24.04. And create/root/gpudrv/drv-audit.sh— it goes through all nodes wherenvidia.com/gpu.present=trueand checks whether that node's combination is in this table; if not, it prints one line<노드> MISSING <태그>(the placeholders are the node and the tag) for each and ends with exit code 1, and if none, it ends with the one lineOKand 0. After creating it, run it and save the output to/root/gpudrv/out/audit.txt(for now it must beOK). - With
/root/gpudrv/k8s/node3.yaml, put a new nodelab-node-3into the cluster. The labels arenvidia.com/gpu.present=true, driver550/90/07, kernel6.8.0-52-generic,ubuntu24.04, and it is a fake node with thekwok.x-k8s.io/node: fakeannotation and a Ready condition (see the example for the format). After applying it, rundrv-audit.shand save the output to/root/gpudrv/out/skew.txt— the new node should be caught. Then, supposing you built that combination and pushed it to the registry, add one line tosupport-matrix.csv, and save the output of running it again to/root/gpudrv/out/skew-fixed.txt(now it must beOK). - Bring up two Pods in the namespace
gpu-drvyou created in step 1. In/root/gpudrv/k8s/cuda12-job.yaml, write the Podcuda12-job— container nametrainer, imagenvcr.io/nvidia/pytorch:24.07-py3, and two conditions inside one term of the required nodeAffinity:nvidia.com/cuda.driver.majorIn550, andfeature.node.kubernetes.io/system-os_release.VERSION_IDIn22.04. In/root/gpudrv/k8s/cuda13-job.yaml, write the Podcuda13-job— the image isnvcr.io/nvidia/pytorch:25.03-py3, and the condition is justnvidia.com/cuda.driver.majorGt560. After applying both, confirm thatcuda12-jobcomes up on lab-node-0 andcuda13-jobwaits, and save thePodScheduledcondition message ofcuda13-jobto/root/gpudrv/out/cuda.txt. - In
/root/gpudrv/k8s/trainer.yaml, write a Deploymenttrainer— namespacegpu-drv,replicas: 4, the Pod label and selectorapp: trainer, containertrainer, imagenvcr.io/nvidia/pytorch:24.07-py3, and with a required podAntiAffinity, make it so that, fortopologyKey: kubernetes.io/hostname, two Pods of the sameapp: trainercannot be on one node (they spread one per node across the four nodes). In/root/gpudrv/k8s/pdb.yaml, write a PodDisruptionBudgettrainer-pdb—minAvailable: 4, with the selectorapp: trainer. After all four Pods are Running, runkubectl drain lab-node-1 --ignore-daemonsets --delete-emptydir-data --timeout=20sand save the output, including standard error, to/root/gpudrv/out/drain-blocked.txt. Being rejected is normal. - Walk the upgrade order as it is. (1) If you raise lab-node-1 to driver
550.90.07, the tag needed is550-5.15.0-107-generic-ubuntu22.04— add that combination tosupport-matrix.csvfirst (you have in effect secured the image). (2) Create room in the budget — lower theminAvailableoftrainer-pdbto3. (3) Run the same drain command again to make it succeed this time, and save the output to/root/gpudrv/out/drain-ok.txt. (4) Change the driver label set of lab-node-1 to550/90/07(you have in effect newly loaded the driver — in this environment no actual installation is done). (5) Return the node withkubectl uncordon lab-node-1. Finally, rundrv-audit.shagain and confirm thatOKcomes out. - The GPU Operator's upgrade controller shows progress with the node label
nvidia.com/gpu-driver-upgrade-state. In/root/gpudrv/out/upgrade-states.txt, write those states in the order they appear in the documentation as eight lines —upgrade-required,cordon-required,pod-deletion-required,drain-required,pod-restart-required,validation-required,uncordon-required, andupgrade-done. And attach the labelnvidia.com/gpu-driver-upgrade-state=upgrade-doneto lab-node-1. Finally, write five lines in/root/gpudrv/out/upgrade-report.txt—NODES=<gpu.present 가 true 인 노드 수>,SKEW=<drv-audit.sh 가 낸 MISSING 줄 수>,LAB_NODE_1_TAG=<drv-tag.sh 가 lab-node-1 에 대해 내는 태그>,MATRIX_ROWS=<support-matrix.csv 의 머리글을 뺀 줄 수>, andUPGRADE_STATE=upgrade-done(the placeholders are the number of nodes where gpu.present is true, the number of MISSING lines drv-audit.sh printed, the tag drv-tag.sh prints for lab-node-1, and the number of lines in support-matrix.csv excluding the header). Extract the numbers and tags with commands from the current state to fill them in.
Notes
- Start with
export KUBECONFIG=/root/.kube/config. You start with three nodes, lab-node-0/1/2, and add one yourself in step 4. Deliverables go in/root/gpudrv, and objects in the namespacegpu-drv. - The driver is not actually installed. This environment has no GPU, no kernel module, and no nvidia-smi. So kernel version and driver version are expressed as node labels, and an upgrade is represented by changing those labels. In a real cluster too, what the scheduler and the Operator look at is, in the end, those labels.
- On the other hand, adding nodes, nodeAffinity, podAntiAffinity, PodDisruptionBudget, eviction refusal, drain, and uncordon are done by the real control plane and so work as they are. What this lab judges is also that side.
- Common mistake: running drain without
--timeout. If it is blocked by the budget, it retries forever. - Common mistake: forgetting uncordon after the drain is finished. That node quietly sits idle without any error.
- Common mistake: writing node names or tags inside the checker. When nodes grow, that tool lies from that day on.
- GPU Driver Upgrades · Precompiled Driver Containers · Safely Drain a Node · Specifying a Disruption Budget
Set up kernel version and driver version as node facts
Work in /root/gpudrv (export KUBECONFIG=/root/.kube/config). First create the namespace gpu-drv (the Pods in later steps come up here). Attach labels to the three nodes. lab-node-0: nvidia.com/gpu.present=true, nvidia.com/cuda.driver.major=550, nvidia.com/cuda.driver.minor=90, nvidia.com/cuda.driver.rev=07, feature.node.kubernetes.io/kernel-version.full=5.15.0-119-generic, feature.node.kubernetes.io/system-os_release.ID=ubuntu, and feature.node.kubernetes.io/system-os_release.VERSION_ID=22.04. lab-node-1: with the same keys, driver 535/183/06, kernel 5.15.0-107-generic, ubuntu 22.04. lab-node-2: driver 550/90/07, kernel 6.8.0-45-generic, ubuntu 24.04. All three nodes have nvidia.com/gpu.present=true.
This environment has neither a GPU nor a driver, so labels are the place of facts. In a real cluster, the kernel and OS labels are attached by nfd-worker and nvidia.com/cuda.driver.* by gpu-feature-discovery. The driver version is split into three label pieces, major.minor.rev — for 550.90.07 they are 550, 90, and 07. You can attach several at once with kubectl label node <이름> <키>=<값> --overwrite (the placeholders are the name, the key, and the value). A value with a dot such as 22.04 does not violate the label value convention.
Compute the driver image tag one node needs
Create /root/gpudrv/drv-tag.sh <노드이름> (the placeholder is the node name). It reads that node's labels and prints the tag of a precompiled driver image as one line. The format is the official documentation's <드라이버브랜치>-<커널판>-<OS태그> (the placeholders are the driver branch, the kernel version, and the OS tag), where the branch is nvidia.com/cuda.driver.major, the kernel version is feature.node.kubernetes.io/kernel-version.full, and the OS tag is system-os_release.ID and system-os_release.VERSION_ID written together (for example, ubuntu and 22.04 give ubuntu22.04). Run it on the three nodes in turn and write three lines in /root/gpudrv/out/tags.txt in the form <노드> <태그> (the placeholders are the node and the tag). Do not write node names or tags inside the script — the grader calls it directly for each node.
The example tag in the NVIDIA documentation is 525-5.15.0-69-generic-ubuntu22.04. The kernel version comes after the branch, and the OS tag after that. The kernel version has hyphens inside it too, so if you try to split the tag from the back to read it, it gets confusing — when making it, just take each piece from its label and join them. If even one label is missing, the tag quietly goes strange. It is better to end with an error when you meet an empty value. If you use // when extracting labels with jq, a missing label and an empty value cannot be told apart.
Cross-check against the combinations in the registry to find mismatched nodes
Create /root/gpudrv/support-matrix.csv. The first line is the header driver_branch,kernel,os_tag, followed by three lines of combinations actually built in the company registry — 550,5.15.0-119-generic,ubuntu22.04, 535,5.15.0-107-generic,ubuntu22.04, and 550,6.8.0-45-generic,ubuntu24.04. And create /root/gpudrv/drv-audit.sh — it goes through all nodes where nvidia.com/gpu.present=true and checks whether that node's combination is in this table; if not, it prints one line <노드> MISSING <태그> (the placeholders are the node and the tag) for each and ends with exit code 1, and if none, it ends with the one line OK and 0. After creating it, run it and save the output to /root/gpudrv/out/audit.txt (for now it must be OK).
This table is precisely "the list of driver images we have". In practice it is the tag list of the NGC registry or the list of images built and pushed in-house, and either way, a tag that does not exist shows up only as a Pod dying with ImagePullBackOff. So the order is to do this cross-check first, before raising the kernel. Fetch the node list as you go — the grader briefly changes one node's kernel label and then calls it. You do not need to build the tag again. Just call the script you made in step 2.
When a node is added, the driver image is the first thing that is short
With /root/gpudrv/k8s/node3.yaml, put a new node lab-node-3 into the cluster. The labels are nvidia.com/gpu.present=true, driver 550/90/07, kernel 6.8.0-52-generic, ubuntu 24.04, and it is a fake node with the kwok.x-k8s.io/node: fake annotation and a Ready condition (see the example for the format). After applying it, run drv-audit.sh and save the output to /root/gpudrv/out/skew.txt — the new node should be caught. Then, supposing you built that combination and pushed it to the registry, add one line to support-matrix.csv, and save the output of running it again to /root/gpudrv/out/skew-fixed.txt (now it must be OK).
A new node is usually installed from the latest image, so its kernel is ahead of the existing nodes. Even if the driver branch is the same, a different kernel version needs a different image — which is why the tag of a precompiled driver contains the kernel version. If you do not know that adding a node is itself image build work, you wander at the spot where the driver Pod dies for lack of an image only on the new node. In this environment, a node is also just an API object, so you can create it with kubectl apply. When adding a line to the CSV, keep the header order (branch, kernel, OS tag).
CUDA is compatible only forward — write that requirement as a label condition
Bring up two Pods in the namespace gpu-drv you created in step 1. In /root/gpudrv/k8s/cuda12-job.yaml, write the Pod cuda12-job — container name trainer, image nvcr.io/nvidia/pytorch:24.07-py3, and two conditions inside one term of the required nodeAffinity: nvidia.com/cuda.driver.major In 550, and feature.node.kubernetes.io/system-os_release.VERSION_ID In 22.04. In /root/gpudrv/k8s/cuda13-job.yaml, write the Pod cuda13-job — the image is nvcr.io/nvidia/pytorch:25.03-py3, and the condition is just nvidia.com/cuda.driver.major Gt 560. After applying both, confirm that cuda12-job comes up on lab-node-0 and cuda13-job waits, and save the PodScheduled condition message of cuda13-job to /root/gpudrv/out/cuda.txt.
A driver does not know a CUDA runtime that came out later than itself. Conversely, an old CUDA container runs fine on a new driver — meaning compatibility is open only one way. So if you write "this container needs driver version so-and-so or higher" as a node label condition, instead of going to a mismatched node and failing during execution, it is blocked at the scheduling stage. If you put two conditions in one term, both must be satisfied. The reason for waiting is in the PodScheduled message of .status.conditions. Gt reads values as integers, so it can be used only on labels that are integers, like the branch number.
What blocks the upgrade is not the driver but the budget
In /root/gpudrv/k8s/trainer.yaml, write a Deployment trainer — namespace gpu-drv, replicas: 4, the Pod label and selector app: trainer, container trainer, image nvcr.io/nvidia/pytorch:24.07-py3, and with a required podAntiAffinity, make it so that, for topologyKey: kubernetes.io/hostname, two Pods of the same app: trainer cannot be on one node (they spread one per node across the four nodes). In /root/gpudrv/k8s/pdb.yaml, write a PodDisruptionBudget trainer-pdb — minAvailable: 4, with the selector app: trainer. After all four Pods are Running, run kubectl drain lab-node-1 --ignore-daemonsets --delete-emptydir-data --timeout=20s and save the output, including standard error, to /root/gpudrv/out/drain-blocked.txt. Being rejected is normal.
To raise the driver you must first bring down that node's GPU workloads, and bringing them down goes through the eviction API. A PodDisruptionBudget is the mechanism that blocks exactly that API — it refuses if the number currently alive would drop below the budget. So it is common that the real cause of "the driver upgrade stopped on one node" is the budget, not the driver. A drain first cordons the node and then evicts Pods one by one. So even when blocked, the node is already unschedulable. If you do not give --timeout, drain retries forever.
Secure the image first, create room, and walk the procedure to the end
Walk the upgrade order as it is. (1) If you raise lab-node-1 to driver 550.90.07, the tag needed is 550-5.15.0-107-generic-ubuntu22.04 — add that combination to support-matrix.csv first (you have in effect secured the image). (2) Create room in the budget — lower the minAvailable of trainer-pdb to 3. (3) Run the same drain command again to make it succeed this time, and save the output to /root/gpudrv/out/drain-ok.txt. (4) Change the driver label set of lab-node-1 to 550/90/07 (you have in effect newly loaded the driver — in this environment no actual installation is done). (5) Return the node with kubectl uncordon lab-node-1. Finally, run drv-audit.sh again and confirm that OK comes out.
The order is the whole of this step. If you drain before securing the image, you empty the node and wait, and if you drain before creating room in the budget, you are rejected as in the previous step. You can edit a PDB with kubectl patch pdb <이름> --type=merge -p '{"spec":{"minAvailable":N}}' (the placeholder is the name). Right after editing the budget, the controller needs a few seconds to recompute status.disruptionsAllowed. When the drain is finished, that node stays cordoned — if you forget uncordon, that node quietly sits idle.
The upgrade state machine and the closing report
The GPU Operator's upgrade controller shows progress with the node label nvidia.com/gpu-driver-upgrade-state. In /root/gpudrv/out/upgrade-states.txt, write those states in the order they appear in the documentation as eight lines — upgrade-required, cordon-required, pod-deletion-required, drain-required, pod-restart-required, validation-required, uncordon-required, and upgrade-done. And attach the label nvidia.com/gpu-driver-upgrade-state=upgrade-done to lab-node-1. Finally, write five lines in /root/gpudrv/out/upgrade-report.txt — NODES=<gpu.present 가 true 인 노드 수>, SKEW=<drv-audit.sh 가 낸 MISSING 줄 수>, LAB_NODE_1_TAG=<drv-tag.sh 가 lab-node-1 에 대해 내는 태그>, MATRIX_ROWS=<support-matrix.csv 의 머리글을 뺀 줄 수>, and UPGRADE_STATE=upgrade-done (the placeholders are the number of nodes where gpu.present is true, the number of MISSING lines drv-audit.sh printed, the tag drv-tag.sh prints for lab-node-1, and the number of lines in support-matrix.csv excluding the header). Extract the numbers and tags with commands from the current state to fill them in.
The point is not to memorize the state machine but the fact that, when an upgrade stalls, you can tell in one line where it stopped. If you extract this label for every node at once with kubectl get node -l nvidia.com/gpu.present -o jsonpath, you can immediately see that the cause differs between a node that stopped at cordon-required and a node that stopped at pod-deletion-required. The numbers in this step must come not from the memory of earlier steps but from the current cluster and the current files — you added a line to the table in step 7 and the driver version also changed, so the values differ from the earlier steps.