Air-Gapped GPU Driver Installation
Registering the nvidia Runtime With containerd
Goal
You register the nvidia runtime in containerd, validate the TOML syntax, and write a CDI spec and a RuntimeClass. Even without an actual GPU, you can verify the correctness of the settings entirely.
Why it matters
The cause of the report "I installed the driver and nvidia-smi works, but the GPU is not visible in containers" is usually a missing runtime registration. And if you pasted a snippet picked up from the internet and it has no effect, it is mostly a config version mismatch — the section headers change wholesale depending on the containerd major version.
There is one more important judgment. It is not changing default_runtime_name to nvidia. If you change it, even Pods that do not use the GPU go through that runtime, and runtime problems spread to the whole cluster. It is safer to designate only the workloads that need it with a RuntimeClass.
Steps
- Create the
/etc/containerddirectory and copy/opt/fixtures/gpu-airgap/containerd/config.toml.baseto/etc/containerd/config.toml. Also create the/root/toolkitworking directory. - Add an nvidia runtime section to
config.toml. The section header is[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia], and it must containruntime_type = "io.containerd.runc.v2". - Add an options section under it. Under
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options], there must beBinaryName = "/usr/bin/nvidia-container-runtime"andSystemdCgroup = true. - Check the value of
default_runtime_nameand write it to/root/toolkit/default.txtas the following two lines.DEFAULT=runc/WHY=runtimeclass(The default runtime must berunc. Do not change it.) - Read
config.tomlwith a TOML parser and save theversionvalue and the list of registered runtime names to/root/toolkit/parsed.txt. Bothruncandnvidiamust show. - Check that
/etc/cdi/nvidia.yamlexists and containskind: nvidia.com/gpu. If it does not exist, create it. Then save the list of top-level keys to/root/toolkit/cdi-keys.txt. - Write
/root/toolkit/runtimeclass.yaml. It must have the four itemsapiVersion: node.k8s.io/v1,kind: RuntimeClass,metadata.name: nvidia, andhandler: nvidia. - Make
/root/toolkit/report.txtwith the following 5 lines.CONFIG_VERSION=2/RUNTIME=nvidia/RUNTIME_TYPE=io.containerd.runc.v2/SYSTEMD_CGROUP=true/HANDLER=nvidia
Notes
- TOML parsing takes the form
python3 -c "import tomllib;d=tomllib.load(open('/etc/containerd/config.toml','rb'));print(d['version'])". - In a real environment,
nvidia-ctk runtime configure --runtime=containerd --set-as-default=falsedoes steps 2–3 for you. - The
handlerof a RuntimeClass must match exactly theruntimes.<이름>(the placeholder is the name) in config.toml. If it does not match, the Pod comes up withRunContainerError. - Common mistake 1: leaving out the quotation marks in the section header. In TOML, a key that contains a dot must be wrapped in quotation marks.
- Common mistake 2: writing
SystemdCgroupas the string"true". It is a TOML boolean, so it istruewithout quotation marks.
Place the base configuration file
Create the /etc/containerd directory and copy /opt/fixtures/gpu-airgap/containerd/config.toml.base to /etc/containerd/config.toml. Also create the /root/toolkit working directory.
The fixture has an abbreviated base configuration. In a real environment you create it with containerd config default.
Add the nvidia runtime section
Add an nvidia runtime section to config.toml. The section header is [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia], and it must contain runtime_type = "io.containerd.runc.v2".
It takes the form runtimes. under the CRI plugin path of config version 2. The runtime_type value is the key.
Specify the runtime options
Add an options section under it. Under [plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options], there must be BinaryName = "/usr/bin/nvidia-container-runtime" and SystemdCgroup = true.
options is a subtable of the runtime section. You need two keys, the binary path and the cgroup driver.
Check the default runtime
Check the value of default_runtime_name and write it to /root/toolkit/default.txt as the following two lines.
DEFAULT=runc / WHY=runtimeclass
(The default runtime must be runc. Do not change it.)
default_runtime_name must be runc. Also record why.
Validate the TOML syntax
Read config.toml with a TOML parser and save the version value and the list of registered runtime names to /root/toolkit/parsed.txt. Both runc and nvidia must show.
Python 3.11 and later has a TOML parser in the standard library. If parsing succeeds, the syntax is fine.
Place the CDI spec
Check that /etc/cdi/nvidia.yaml exists and contains kind: nvidia.com/gpu. If it does not exist, create it. Then save the list of top-level keys to /root/toolkit/cdi-keys.txt.
It has the same structure as what you made in the earlier course. Here you check only the placement path and the minimum keys.
Write the RuntimeClass
Write /root/toolkit/runtimeclass.yaml. It must have the four items apiVersion: node.k8s.io/v1, kind: RuntimeClass, metadata.name: nvidia, and handler: nvidia.
You need four things: apiVersion, kind, metadata.name, and handler. handler must equal the runtime name in config.toml.
Registration verification report
Make /root/toolkit/report.txt with the following 5 lines.
CONFIG_VERSION=2 / RUNTIME=nvidia / RUNTIME_TYPE=io.containerd.runc.v2 / SYSTEMD_CGROUP=true / HANDLER=nvidia
Get the values by parsing the files you wrote. Write the paths as absolute paths.