Matching driver versions — why the kernel version sits inside the image tag
In one line
What a GPU driver deployed as a container actually does is load a module into the host kernel, so when the kernel changes the driver must be rebuilt too (which is why the tag of a precompiled image contains the kernel version), and to load the driver the node's GPU workloads must first be emptied (which is why a PodDisruptionBudget blocks upgrades).
Why this becomes a problem
An ordinary container runs in an isolated user space. Whatever is installed on the host, it uses the libraries inside the image, so the intuition that "if you make it a container it becomes independent of the host" holds.
A driver is not like that. The NVIDIA documentation gives this as the reason driver upgrades need special consideration. The driver kernel module must be unloaded and loaded again every time the driver container restarts. And the procedure has five steps — stop all clients using the driver, unload the current kernel module, bring up the new driver Pod, load the new module, and turn the clients back on.
Two properties come out of this. First, the driver depends on the host kernel. A kernel module must be compiled to match the kernel version, so if the kernel goes up by one, a driver for that kernel is separately needed. Second, replacing the driver always touches that node's GPU workloads. That is because to unload the module, there must be no process using it.
How it works — the image tag is the contract
A driver container in the default mode fetches the kernel headers and compiler on the node and builds the module at boot. It needs the internet, takes time, and eats CPU. So NVIDIA separately publishes precompiled driver containers. The documentation says the benefit is especially worthwhile where internet access is restricted and where resources are tight.
The tag of a precompiled image looks like this.
<driver-branch>-<linux-kernel-version>-<os-tag>
예: 525-5.15.0-69-generic-ubuntu22.04
That three pieces are baked into one tag is the contract of this method. Even if the driver branch is the same, a different kernel version is a different image, and even if the kernel version is the same, a different OS is yet another image. So in operations, the following all become "image work".
- You added a node but it was installed from a new image and its kernel is ahead → an image for that combination is needed
- A security patch raised the kernel by one → an image for that combination is needed
- You move a node pool from 22.04 to 24.04 → an image for that combination is needed
How does it show up when it does not exist? Only on that node, the driver Pod cannot pull the image and dies. The other nodes are fine, so no cluster-level alert fires, and only that node's GPU quietly drops out. The documentation states clearly that the supported combinations are fixed in a table, and that if you use a kernel variant not on the list, you must build the image yourself and push it to your own registry.
So mature teams keep "the list of driver images we have" as a ledger and periodically cross-check it against the nodes' current combinations. When a kernel upgrade plan is made, they run that cross-check first and build the missing combinations in advance. If you do not keep this order, you end up emptying the node and waiting for an image.
CUDA is compatible only forward
There is a version issue between the driver and the CUDA runtime too, and the direction is open only one way. An old CUDA container runs fine on a new driver. The reverse does not work. That is because a driver does not know a CUDA runtime that came out later than itself.
What this means in practice is simple. When you start using a new framework image, first look at whether the lowest driver version in the cluster can handle that image. And write that requirement into the Pod spec, not in people's memory — if you put a nodeAffinity condition on the nvidia.com/cuda.driver.major label that gpu-feature-discovery attaches, instead of going to a mismatched node and failing during execution, it is blocked at the scheduling stage. Pulling the failure forward is the value of this condition.
Where the upgrade stalls
To raise the driver, you must bring down that node's GPU workloads, and bringing them down goes through the eviction API. And in front of the eviction API there is a PodDisruptionBudget. If the number currently alive would drop below the budget, the API refuses the request.
error when evicting pods/"trainer-..." (will retry after 5s):
Cannot evict pod as it would violate the pod's disruption budget.
kubectl drain first cordons the node and then evicts Pods one by one, so even when blocked, the node is already unschedulable. If you do not notice here, that node is left in an awkward state in which it accepts no new Pods and cannot get the driver raised either. A drain run without --timeout retries forever, so on screen it looks merely "slow".
The GPU Operator's upgrade controller automates this procedure and writes the progress state on the node label nvidia.com/gpu-driver-upgrade-state. The states the documentation defines are these.
| State | Meaning |
|---|---|
upgrade-required |
The driver Pod is not the latest |
cordon-required |
It is time to mark the node unschedulable |
wait-for-jobs-required |
Wait for the specified jobs to finish |
pod-deletion-required |
Delete the Pods that were allocated a GPU |
drain-required |
Deleting Pods is not enough, so drain the node |
pod-restart-required |
Restart the driver Pod to bring up the new version |
validation-required |
Validate the new driver |
uncordon-required |
Return the node to schedulable |
upgrade-done |
Finished |
The value of this label is not the state names themselves but that you can tell in one line where it stopped. A node that stopped at cordon-required and a node that stopped at pod-deletion-required have different causes. If you extract all the nodes at once, you immediately see how many piled up at which stage.
What the documentation warns most strongly about is drain.enable. There is a reason its default is off — a drain drives out from that node even workloads unrelated to the GPU. It says to first adjust the GPU Pod deletion settings, and only when that is not enough to turn on drain, narrowing the scope with podSelector.
What it looks like in the field
First, "only one node can't see the GPU". Either the kernel went up on only that node, or only that node joined later and its image is different. The check ends by putting that node's kernel label and the driver Pod's image tag side by side.
Second, "the upgrade has been stuck on one node for three hours". The reason you never find it however much driver log you read is that the cause is not the driver. Either a PDB is refusing eviction, or the job the controller was told to wait for has not finished. If you look at the upgrade state label first, it is over in 3 seconds.
Third, preinstalled drivers. In teams that bake the driver into the image, the Operator does not manage the driver. The documentation also says that the Operator does not manage the lifecycle of a driver preinstalled on the host. It looks convenient, but in exchange version management becomes entirely a person's job, so which is better depends on the maturity of the team's image pipeline.
Fourth, the honest limit of this environment. The lab does not actually install a driver. There is no GPU, no kernel module, and no nvidia-smi. So kernel version and driver version are expressed as node labels, and an upgrade is represented by changing those labels — in a real cluster too, what the scheduler and the Operator look at is, in the end, those labels. On the other hand, adding nodes, PodDisruptionBudget judgments, eviction refusals, and drain and uncordon are carried out as they are by the real control plane.
Reference documents
- GPU driver upgrades (the upgrade state machine and policy): https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/gpu-driver-upgrades.html
- Precompiled driver containers (tag rules and constraints): https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/precompiled-drivers.html
- The platform support table: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/platform-support.html
- Safely draining a node: https://kubernetes.io/docs/tasks/administer-cluster/safely-drain-node/
- Specifying a PodDisruptionBudget: https://kubernetes.io/docs/tasks/run-application/configure-pdb/
What you will do in the next lab
Taking node labels as the place of facts, you build a tool that computes the driver image tag a node needs and a checker that cross-checks it against the company image list. You add one node and confirm that the checker tells you first that only that node lacks a combination, and you write the CUDA requirement as a nodeAffinity and see a mismatched requirement blocked at the scheduling stage. And after passing the spot where a PodDisruptionBudget actually refuses a drain, you walk the upgrade procedure in order, from securing the image through uncordon, to the end.