TT Lab
Get started
Learn Learning paths Courses

Running Rootless Podman

Reading and Writing a CDI Specification

Continue in TT Lab

In one line

CDI is a specification, written in YAML, of "what has to be done to put this device into a container." It turned the black box of runtime hooks into a standard file.

Why this was needed

In the past, to use a GPU with docker you had to plug in a runtime hook called nvidia-container-runtime. Right before the container started, that hook ran, mounted the driver libraries, and added the device nodes. It worked, but it had problems.

CDI turned this into a declarative specification. If you write in /etc/cdi/*.yaml "when this device name is requested, add these device nodes, these mounts, and these environment variables," a runtime that supports CDI (podman, containerd, CRI-O) carries it out as written.

How it works

Spec structure

cdiVersion: "0.6.0"
kind: nvidia.com/gpu

devices:
  - name: "0"
    containerEdits:
      deviceNodes:
        - path: /dev/nvidia0
        - path: /dev/nvidiactl
        - path: /dev/nvidia-uvm
      mounts:
        - hostPath: /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.550.90.07
          containerPath: /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.550.90.07
          options: ["ro", "nosuid", "nodev", "bind"]
      env:
        - NVIDIA_VISIBLE_DEVICES=0
      hooks:
        - hookName: createContainer
          path: /usr/bin/nvidia-ctk
          args: ["nvidia-ctk", "hook", "update-ldcache"]

  - name: all
    containerEdits:
      deviceNodes:
        - path: /dev/nvidia0
        - path: /dev/nvidiactl

Key rules.

In actual use

# 스펙 생성 (실제 GPU 가 있는 호스트에서)
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml

# 등록된 장치 확인
nvidia-ctk cdi list
# nvidia.com/gpu=0
# nvidia.com/gpu=all

# 컨테이너에서 사용
podman run --rm --device nvidia.com/gpu=all nvidia/cuda:12.4.1-base nvidia-smi

The counterpart of docker's --gpus all is --device nvidia.com/gpu=all. Recent podman also accepts --gpus as a compatibility flag, but the CDI notation is the official one.

GPUs work in rootless mode too. The GPU driver runs in kernel space, so the container's privilege level has nothing to do with performance. If it can read the system-wide /etc/cdi/nvidia.yaml, it uses that as it is, and if not, you generate the spec in user space and point to that directory.

Registration on the containerd side

In Kubernetes, containerd reads CDI. You need an enabling setting in config.toml, and the older way of registering a runtime separately is still in use.

[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia]
  runtime_type = "io.containerd.runc.v2"

[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia.options]
  BinaryName = "/usr/bin/nvidia-container-runtime"
  SystemdCgroup = true

And a workload that uses that runtime is selected with a RuntimeClass.

apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
  name: nvidia
handler: nvidia

The trick is not to change default_runtime_name to nvidia. If you do, even Pods that do not use a GPU go through that runtime. It is safer to specify only the workloads that need it with a RuntimeClass.

What it looks like in the field

The GPU is not detected after a driver update. Nine times out of ten, you forgot to regenerate the CDI spec (nvidia-ctk cdi generate). The spec contains paths with the version embedded for the library files, so when the driver is upgraded those paths disappear.

nvidia-smi works inside the container but the CUDA program fails. Only some of the required libraries were mounted. You have to check the mounts list of the spec.

What you will do in the next lab

You write a CDI spec directly in /etc/cdi/. There is no real GPU, so grading covers the spec's syntax and placement, and the format of the --device argument. Those are exactly the places where mistakes happen in the field.