Reading and Writing a CDI Specification
In one line
CDI is a specification, written in YAML, of "what has to be done to put this device into a container." It turned the black box of runtime hooks into a standard file.
Why this was needed
In the past, to use a GPU with docker you had to plug in a runtime hook called nvidia-container-runtime. Right before the container started, that hook ran, mounted the driver libraries, and added the device nodes. It worked, but it had problems.
- You cannot tell from outside what it does. The logic inside the hook is a black box.
- Every vendor builds its own hook. NVIDIA, AMD, FPGA, and InfiniBand each work differently.
- You have to replace the runtime. You had to change the configuration so that nvidia-container-runtime is used instead of runc.
CDI turned this into a declarative specification. If you write in /etc/cdi/*.yaml "when this device name is requested, add these device nodes, these mounts, and these environment variables," a runtime that supports CDI (podman, containerd, CRI-O) carries it out as written.
How it works
Spec structure
cdiVersion: "0.6.0"
kind: nvidia.com/gpu
devices:
- name: "0"
containerEdits:
deviceNodes:
- path: /dev/nvidia0
- path: /dev/nvidiactl
- path: /dev/nvidia-uvm
mounts:
- hostPath: /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.550.90.07
containerPath: /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.550.90.07
options: ["ro", "nosuid", "nodev", "bind"]
env:
- NVIDIA_VISIBLE_DEVICES=0
hooks:
- hookName: createContainer
path: /usr/bin/nvidia-ctk
args: ["nvidia-ctk", "hook", "update-ldcache"]
- name: all
containerEdits:
deviceNodes:
- path: /dev/nvidia0
- path: /dev/nvidiactl
Key rules.
kindhas the form<벤더>/<클래스>(the placeholders are the vendor and the class).nvidia.com/gpu,amd.com/gpu,intel.com/fpga. The vendor part must be in domain form.- A device name is referenced as
kind=name.nvidia.com/gpu=0,nvidia.com/gpu=all. containerEditsis the content to add to the OCI spec when that device is attached. The four typical ones are deviceNodes, mounts, env, and hooks.- Put spec files in
/etc/cdior/var/run/cdi. For rootless, you can also specify a user path.
In actual use
# 스펙 생성 (실제 GPU 가 있는 호스트에서)
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml
# 등록된 장치 확인
nvidia-ctk cdi list
# nvidia.com/gpu=0
# nvidia.com/gpu=all
# 컨테이너에서 사용
podman run --rm --device nvidia.com/gpu=all nvidia/cuda:12.4.1-base nvidia-smi
The counterpart of docker's --gpus all is --device nvidia.com/gpu=all. Recent podman also accepts --gpus as a compatibility flag, but the CDI notation is the official one.
GPUs work in rootless mode too. The GPU driver runs in kernel space, so the container's privilege level has nothing to do with performance. If it can read the system-wide /etc/cdi/nvidia.yaml, it uses that as it is, and if not, you generate the spec in user space and point to that directory.
Registration on the containerd side
In Kubernetes, containerd reads CDI. You need an enabling setting in config.toml, and the older way of registering a runtime separately is still in use.
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia]
runtime_type = "io.containerd.runc.v2"
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options]
BinaryName = "/usr/bin/nvidia-container-runtime"
SystemdCgroup = true
And a workload that uses that runtime is selected with a RuntimeClass.
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: nvidia
handler: nvidia
The trick is not to change default_runtime_name to nvidia. If you do, even Pods that do not use a GPU go through that runtime. It is safer to specify only the workloads that need it with a RuntimeClass.
What it looks like in the field
The GPU is not detected after a driver update. Nine times out of ten, you forgot to regenerate the CDI spec (nvidia-ctk cdi generate). The spec contains paths with the version embedded for the library files, so when the driver is upgraded those paths disappear.
nvidia-smi works inside the container but the CUDA program fails. Only some of the required libraries were mounted. You have to check the mounts list of the spec.
What you will do in the next lab
You write a CDI spec directly in /etc/cdi/. There is no real GPU, so grading covers the spec's syntax and placement, and the format of the --device argument. Those are exactly the places where mistakes happen in the field.