TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

Reproducing the GPU Incident — From Config Merge to Stalled Rollout

Continue in TT Lab

Goal

You walk through, in eight steps, the accident that made a GPU cluster unusable all at once. You build by hand the spot where the containerd configuration gets merged, the habit of checking what is loaded, the two gates of resource advertisement and runtime handler, the arithmetic of time-slicing, and even the shape of a cluster whose rollout has stalled.

Why it matters

What catches people in GPU operations is not the driver but the moment they believe the configuration was applied. Even if the drop-in file has the nvidia runtime written in it intact, the loaded runtime can be just runc; even if the toolkit Pod is Ready, the handler is not registered if containerd was not restarted; and already running Pods say nothing even when all of this is broken. That is why a failure stays latent and bursts all at once at a node reboot. So this lab makes you look only at "what was actually loaded and what was actually scheduled", not "what was written". That one habit shortens hours of cause-finding to 3 seconds.

This environment has no real GPU, no containerd, and no GPU Operator. Steps 1–4 are handled with configuration files and a TOML parser, and steps 5–8 are handled on the real control plane brought up by kwok and 3 fake nodes (lab-node-0/1/2). Scheduling, resource advertisement, and DaemonSet rollouts are handled by real controllers and so are reproduced as they are, but containers do not actually run, and a GPU plugged into a node is just a number on the node object.

Steps

  1. Create /root/gpuop/etc/containerd/config.toml. It must have version = 2 and an imports array at the top level, and one entry of imports must end with conf.d/*.toml. Put default_runtime_name = "runc" under plugins."io.containerd.grpc.v1.cri".containerd, and in the plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc table put runtime_type = "io.containerd.runc.v2" and options.SystemdCgroup = true.
  2. Create /root/gpuop/etc/containerd/conf.d/99-nvidia.toml. In the plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia table put runtime_type = "io.containerd.runc.v2", and under it, in options, put BinaryName = "/usr/bin/nvidia-container-runtime" and SystemdCgroup = true.
  3. Read the two files with tomllib and write the judged result in /root/gpuop/out/merge.txt as five lines. MAIN_RUNTIMES= is the runtime names the main configuration defines directly, DROPIN_RUNTIMES= the names the drop-in defines, CONFLICT= is yes if both sides define the same table, CONFLICT_TABLE= is the path of the colliding table (plugins."io.containerd.grpc.v1.cri".containerd.runtimes), and the last is LOADED_RUNTIMES=. Only the last line is not a computed value but an observed value — the runtime handler loaded on that node on the day of the accident, as confirmed with containerd config dump, was only runc.
  4. Create /root/gpuop/bin/check-runtime.sh. It must call containerd config dump without an absolute path, look for containerd.runtimes.nvidia in that output, and end with 0 if found and 1 if not. It must not read the /etc/containerd/config.toml path — the grader flags that as a wrong answer.
  5. In /root/gpuop/k8s/runtimeclass.yaml, write a RuntimeClass whose name and handler are both nvidia, and in /root/gpuop/k8s/gpu-pod.yaml, write a Pod with the namespace gpu-lab, the name cuda-probe, runtimeClassName: nvidia, and nvidia.com/gpu: 1 in the container resource limits. Create the namespace first and apply both. The Pod does not come up — save the reason and message of the PodScheduled condition as two lines in /root/gpuop/out/unscheduled.txt.
  6. With kubectl patch node lab-node-0 --subresource=status, put nvidia.com/gpu as "4" in both capacity and allocatable. When the Pod becomes Running, save the names and advertised amounts of the three nodes as a two-column table in /root/gpuop/out/node-gpu.txt.
  7. In /root/gpuop/k8s/time-slicing.yaml, write a ConfigMap with the namespace gpu-operator and the name time-slicing-config. data.any is a YAML string, and inside it there must be sharing.timeSlicing.failRequestsGreaterThanOne: true and the first item of the resources array must be name: nvidia.com/gpu with replicas: 5. After applying it, raise the advertised amount of lab-node-1 to "20", and write seven lines in /root/gpuop/out/slots.txt — PHYSICAL_GPUS REPLICAS ADVERTISED VRAM_PER_GPU_GIB VRAM_GUARANTEED_PER_SLOT_GIB VRAM_SHARED_PER_GPU_GIB ISOLATION. The equipment is assumed to be 4 A100s of 40GiB.
  8. In /root/gpuop/k8s/toolkit-ds.yaml, write and apply a DaemonSet with the namespace gpu-operator, the name nvidia-container-toolkit, updateStrategy.rollingUpdate.maxUnavailable: 1, the image nvcr.io/nvidia/k8s/container-toolkit:v1.16.2, and a memory request of 128Mi. When all three Pods are Ready, raise the image to v1.17.0 and change the memory request to 64Gi at the same time. When the rollout stalls, write six lines in /root/gpuop/out/rollout.txt — DESIRED UPDATED READY OLD_IMAGE_PODS MAX_UNAVAILABLE, and BLOCKER=insufficient-memory.

Notes

Reproduce the accident node's main containerd configuration

Create /root/gpuop/etc/containerd/config.toml. It must have version = 2 and an imports array at the top level, and one entry of imports must end with conf.d/*.toml. Put default_runtime_name = "runc" under plugins."io.containerd.grpc.v1.cri".containerd, and in the plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc table put runtime_type = "io.containerd.runc.v2" and options.SystemdCgroup = true.

TOML builds its hierarchy by table names, not by indentation. A key that contains dots must be wrapped in double quotation marks to be read as one unit. After writing it, do not check by eye; read it once with a parser.

Reproduce the drop-in the toolkit drops off

Create /root/gpuop/etc/containerd/conf.d/99-nvidia.toml. In the plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia table put runtime_type = "io.containerd.runc.v2", and under it, in options, put BinaryName = "/usr/bin/nvidia-container-runtime" and SystemdCgroup = true.

A drop-in declares one more runtime of its own under the same table path as the main configuration. The key is the binary name in the options — that wrapper inserts the devices and libraries right before the container is created. The cgroup driver must not disagree with the main configuration side.

Judge whether the two files collide over the same table

Read the two files with tomllib and write the judged result in /root/gpuop/out/merge.txt as five lines. MAIN_RUNTIMES= is the runtime names the main configuration defines directly, DROPIN_RUNTIMES= the names the drop-in defines, CONFLICT= is yes if both sides define the same table, CONFLICT_TABLE= is the path of the colliding table (plugins."io.containerd.grpc.v1.cri".containerd.runtimes), and the last is LOADED_RUNTIMES=. Only the last line is not a computed value but an observed value — the runtime handler loaded on that node on the day of the accident, as confirmed with containerd config dump, was only runc.

Leave the judgment not to your hands but to the parser. Read the import list of the main configuration, expand the glob, and gather the names defined in the runtime tables on both sides. Only the last line is not a computation but an observed value from the day of the accident, and that value is written in the instructions.

A check script that looks at the loaded configuration

Create /root/gpuop/bin/check-runtime.sh. It must call containerd config dump without an absolute path, look for containerd.runtimes.nvidia in that output, and end with 0 if found and 1 if not. It must not read the /etc/containerd/config.toml path — the grader flags that as a wrong answer.

The basis of judgment is not the file but the configuration held by the running daemon. Call the command without an absolute path — the grader sets up a fake daemon at the front and actually runs your script twice. It must end with a nonzero value when it finds nothing for the check to do its job.

The RuntimeClass, the GPU Pod, and Pending

In /root/gpuop/k8s/runtimeclass.yaml, write a RuntimeClass whose name and handler are both nvidia, and in /root/gpuop/k8s/gpu-pod.yaml, write a Pod with the namespace gpu-lab, the name cuda-probe, runtimeClassName: nvidia, and nvidia.com/gpu: 1 in the container resource limits. Create the namespace first and apply both. The Pod does not come up — save the reason and message of the PodScheduled condition as two lines in /root/gpuop/out/unscheduled.txt.

A RuntimeClass is cluster-scoped so it has no namespace, and the name and the handler are different fields. It is normal for the Pod not to come up — leave the reason written in the scheduling condition in a file. It takes a few seconds for the condition to be filled in, so instead of sleeping for a fixed time, wait while checking the condition.

What the device plugin does, through the node status

With kubectl patch node lab-node-0 --subresource=status, put nvidia.com/gpu as "4" in both capacity and allocatable. When the Pod becomes Running, save the names and advertised amounts of the three nodes as a two-column table in /root/gpuop/out/node-gpu.txt.

Extended resources are under the status of the node object, and status is a separate subresource, so a plain patch does not take effect. You must fill in both capacity and allocatable, and write the value as a string. Once a slot appears, the scheduler picks up the waiting Pod again.

Read the time-slicing configuration and compute slots and memory

In /root/gpuop/k8s/time-slicing.yaml, write a ConfigMap with the namespace gpu-operator and the name time-slicing-config. data.any is a YAML string, and inside it there must be sharing.timeSlicing.failRequestsGreaterThanOne: true and the first item of the resources array must be name: nvidia.com/gpu with replicas: 5. After applying it, raise the advertised amount of lab-node-1 to "20", and write seven lines in /root/gpuop/out/slots.txt — PHYSICAL_GPUS REPLICAS ADVERTISED VRAM_PER_GPU_GIB VRAM_GUARANTEED_PER_SLOT_GIB VRAM_SHARED_PER_GPU_GIB ISOLATION. The equipment is assumed to be 4 A100s of 40GiB.

The data values of a ConfigMap must be strings, so if you write the configuration as a mapping, the apply is rejected. Put it in as a block scalar. The trap in the computation is the memory line — slots becoming five times as many does not mean the memory is divided into five.

Stall the rollout and report the split state

In /root/gpuop/k8s/toolkit-ds.yaml, write and apply a DaemonSet with the namespace gpu-operator, the name nvidia-container-toolkit, updateStrategy.rollingUpdate.maxUnavailable: 1, the image nvcr.io/nvidia/k8s/container-toolkit:v1.16.2, and a memory request of 128Mi. When all three Pods are Ready, raise the image to v1.17.0 and change the memory request to 64Gi at the same time. When the rollout stalls, write six lines in /root/gpuop/out/rollout.txt — DESIRED UPDATED READY OLD_IMAGE_PODS MAX_UNAVAILABLE, and BLOCKER=insufficient-memory.

First wait until all three nodes are ready, and only then push the new spec. Otherwise you cannot tell what stalled. The node capacity is 32Gi, so if you request more memory than that, no node can accept that Pod. Read the numbers in the report straight from the DaemonSet status and write them.