Air-Gapped GPU Driver Installation
Registering the Runtime With containerd
In one line
To use the GPU in a container, all three of the driver (kernel), the toolkit (the injection tool), and the runtime registration (config.toml) must be present. If even one is missing, the symptom is the same: "the GPU is not visible".
Why this was needed
You installed the driver and nvidia-smi works fine, but inside a container the GPU is not visible. It is a common situation, and the cause is usually that the runtime registration was not done.
How it works
Three layers
| Layer | What | How to check |
|---|---|---|
| Kernel | The nvidia.ko module, the /dev/nvidia* device nodes |
lsmod | grep nvidia, ls /dev/nvidia* |
| User space | libcuda.so, libnvidia-ml.so, and so on |
nvidia-smi, ldconfig -p | grep nvidia |
| Container injection | nvidia-container-cli, the CDI spec, runtime registration |
nvidia-ctk cdi list, config.toml |
There is no kernel inside a container (it is shared with the host). So what is needed is putting the user-space libraries and device nodes into the container, and the toolkit does that.
containerd config.toml
You create the default file with containerd config default and then add the runtime. Do not assemble the section headers by hand — the section headers change wholesale depending on the containerd major version.
config version 2 (containerd 1.x):
version = 2
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia]
runtime_type = "io.containerd.runc.v2"
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options]
BinaryName = "/usr/bin/nvidia-container-runtime"
SystemdCgroup = true
In config version 3 (containerd 2.x), the CRI settings are split into runtime and image.
version = 3
[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.nvidia.options]
BinaryName = '/usr/bin/nvidia-container-runtime'
SystemdCgroup = true
A version 2 file is still supported in 2.x and is converted automatically. Version 1 is not supported from 2.0. If you pasted a snippet picked up from the internet and it has no effect, it is mostly this version mismatch.
nvidia-ctk does this editing for you.
nvidia-ctk runtime configure --runtime=containerd --set-as-default=false
--set-as-default=false matters. If you change the default runtime to nvidia, even Pods that do not use the GPU go through that runtime. If that runtime has a problem, the whole cluster is affected.
SystemdCgroup
The default is false. But on systemd-based hosts true is recommended. The reason is that if you change only kubelet to the systemd cgroup driver and leave containerd as it is, the two managers end up with different views of cgroups. This mismatch shows up as resource limits behaving strangely or Pods being evicted differently from expectations.
RuntimeClass
It lets workloads choose the runtime you registered.
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: nvidia
handler: nvidia # config.toml 의 runtimes.<이름> 과 일치해야 한다
spec:
runtimeClassName: nvidia
containers:
- name: cuda
resources:
limits:
nvidia.com/gpu: 1
The handler value and the runtime name in config.toml must match exactly. If they do not match, the Pod comes up with RunContainerError and the message says "no runtime for ... is configured".
The trap of managed nodes
On managed node groups, config.toml is often regenerated at node bootstrap. A setting fixed by hand disappears the moment the node is replaced, and, worse, it remains on only some nodes and creates differences that cannot be reproduced. If you need to change the settings, you have to fix the bootstrap script, the launch template, or the node image itself. What you fixed by going into a node is a diagnosis, not a deployment.
What it looks like in the field
The GPU is not picked up after a driver update. You forgot to regenerate the CDI spec. The spec contains library paths with versions baked in.
Only GPU Pods start slowly. It is the cost of injecting libraries and refreshing the ldcache at every container start. It stands out more if the image is large.
What you will do in the next lab
You register the nvidia runtime in config.toml and verify it with a TOML parser. You place the CDI spec and write the RuntimeClass YAML. Even without an actual GPU, you can verify the correctness of the settings entirely.