Reproducing the GPU Incident — From Config Merge to Stalled Rollout
Goal
You walk through, in eight steps, the accident that made a GPU cluster unusable all at once. You build by hand the spot where the containerd configuration gets merged, the habit of checking what is loaded, the two gates of resource advertisement and runtime handler, the arithmetic of time-slicing, and even the shape of a cluster whose rollout has stalled.
Why it matters
What catches people in GPU operations is not the driver but the moment they believe the configuration was applied. Even if the drop-in file has the nvidia runtime written in it intact, the loaded runtime can be just runc; even if the toolkit Pod is Ready, the handler is not registered if containerd was not restarted; and already running Pods say nothing even when all of this is broken. That is why a failure stays latent and bursts all at once at a node reboot. So this lab makes you look only at "what was actually loaded and what was actually scheduled", not "what was written". That one habit shortens hours of cause-finding to 3 seconds.
This environment has no real GPU, no containerd, and no GPU Operator. Steps 1–4 are handled with configuration files and a TOML parser, and steps 5–8 are handled on the real control plane brought up by kwok and 3 fake nodes (lab-node-0/1/2). Scheduling, resource advertisement, and DaemonSet rollouts are handled by real controllers and so are reproduced as they are, but containers do not actually run, and a GPU plugged into a node is just a number on the node object.
Steps
- Create
/root/gpuop/etc/containerd/config.toml. It must haveversion = 2and animportsarray at the top level, and one entry ofimportsmust end withconf.d/*.toml. Putdefault_runtime_name = "runc"underplugins."io.containerd.grpc.v1.cri".containerd, and in theplugins."io.containerd.grpc.v1.cri".containerd.runtimes.runctable putruntime_type = "io.containerd.runc.v2"andoptions.SystemdCgroup = true. - Create
/root/gpuop/etc/containerd/conf.d/99-nvidia.toml. In theplugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidiatable putruntime_type = "io.containerd.runc.v2", and under it, inoptions, putBinaryName = "/usr/bin/nvidia-container-runtime"andSystemdCgroup = true. - Read the two files with
tomlliband write the judged result in/root/gpuop/out/merge.txtas five lines.MAIN_RUNTIMES=is the runtime names the main configuration defines directly,DROPIN_RUNTIMES=the names the drop-in defines,CONFLICT=isyesif both sides define the same table,CONFLICT_TABLE=is the path of the colliding table (plugins."io.containerd.grpc.v1.cri".containerd.runtimes), and the last isLOADED_RUNTIMES=. Only the last line is not a computed value but an observed value — the runtime handler loaded on that node on the day of the accident, as confirmed withcontainerd config dump, was onlyrunc. - Create
/root/gpuop/bin/check-runtime.sh. It must callcontainerd config dumpwithout an absolute path, look forcontainerd.runtimes.nvidiain that output, and end with0if found and1if not. It must not read the/etc/containerd/config.tomlpath — the grader flags that as a wrong answer. - In
/root/gpuop/k8s/runtimeclass.yaml, write a RuntimeClass whose name andhandlerare bothnvidia, and in/root/gpuop/k8s/gpu-pod.yaml, write a Pod with the namespacegpu-lab, the namecuda-probe,runtimeClassName: nvidia, andnvidia.com/gpu: 1in the container resourcelimits. Create the namespace first and apply both. The Pod does not come up — save thereasonandmessageof thePodScheduledcondition as two lines in/root/gpuop/out/unscheduled.txt. - With
kubectl patch node lab-node-0 --subresource=status, putnvidia.com/gpuas"4"in bothcapacityandallocatable. When the Pod becomesRunning, save the names and advertised amounts of the three nodes as a two-column table in/root/gpuop/out/node-gpu.txt. - In
/root/gpuop/k8s/time-slicing.yaml, write a ConfigMap with the namespacegpu-operatorand the nametime-slicing-config.data.anyis a YAML string, and inside it there must besharing.timeSlicing.failRequestsGreaterThanOne: trueand the first item of theresourcesarray must bename: nvidia.com/gpuwithreplicas: 5. After applying it, raise the advertised amount oflab-node-1to"20", and write seven lines in/root/gpuop/out/slots.txt—PHYSICAL_GPUSREPLICASADVERTISEDVRAM_PER_GPU_GIBVRAM_GUARANTEED_PER_SLOT_GIBVRAM_SHARED_PER_GPU_GIBISOLATION. The equipment is assumed to be 4 A100s of 40GiB. - In
/root/gpuop/k8s/toolkit-ds.yaml, write and apply a DaemonSet with the namespacegpu-operator, the namenvidia-container-toolkit,updateStrategy.rollingUpdate.maxUnavailable: 1, the imagenvcr.io/nvidia/k8s/container-toolkit:v1.16.2, and a memory request of128Mi. When all three Pods are Ready, raise the image tov1.17.0and change the memory request to64Giat the same time. When the rollout stalls, write six lines in/root/gpuop/out/rollout.txt—DESIREDUPDATEDREADYOLD_IMAGE_PODSMAX_UNAVAILABLE, andBLOCKER=insufficient-memory.
Notes
- Grading of step 4 does not end at a static check. A fake
containerdis set up at the front of PATH and your script is actually run twice — it passes only if it fails when given a dump without nvidia and succeeds when given a dump with it. kubectl patch --subresource=statusworks with kubectl 1.24 or later. The node capacity is CPU 8 and memory 32Gi for all three nodes.- Common mistake 1: writing the TOML table name without quotation marks, like
[plugins.io.containerd.grpc.v1.cri...]. A key that contains dots must be wrapped in double quotation marks to be read as one unit; otherwise the parser builds an entirely different hierarchy. - Common mistake 2: in step 7, writing
data.anyas a YAML mapping. Thedatavalues of a ConfigMap must be strings, so you must put it in as a|-block scalar. - Common mistake 3: in step 8, pushing the new spec before the three Pods are Ready. If you push it from a state that was not ready from the start, you cannot tell what it stalled on.
Reproduce the accident node's main containerd configuration
Create /root/gpuop/etc/containerd/config.toml. It must have version = 2 and an imports array at the top level, and one entry of imports must end with conf.d/*.toml. Put default_runtime_name = "runc" under plugins."io.containerd.grpc.v1.cri".containerd, and in the plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runc table put runtime_type = "io.containerd.runc.v2" and options.SystemdCgroup = true.
TOML builds its hierarchy by table names, not by indentation. A key that contains dots must be wrapped in double quotation marks to be read as one unit. After writing it, do not check by eye; read it once with a parser.
Reproduce the drop-in the toolkit drops off
Create /root/gpuop/etc/containerd/conf.d/99-nvidia.toml. In the plugins."io.containerd.grpc.v1.cri".containerd.runtimes.nvidia table put runtime_type = "io.containerd.runc.v2", and under it, in options, put BinaryName = "/usr/bin/nvidia-container-runtime" and SystemdCgroup = true.
A drop-in declares one more runtime of its own under the same table path as the main configuration. The key is the binary name in the options — that wrapper inserts the devices and libraries right before the container is created. The cgroup driver must not disagree with the main configuration side.
Judge whether the two files collide over the same table
Read the two files with tomllib and write the judged result in /root/gpuop/out/merge.txt as five lines. MAIN_RUNTIMES= is the runtime names the main configuration defines directly, DROPIN_RUNTIMES= the names the drop-in defines, CONFLICT= is yes if both sides define the same table, CONFLICT_TABLE= is the path of the colliding table (plugins."io.containerd.grpc.v1.cri".containerd.runtimes), and the last is LOADED_RUNTIMES=. Only the last line is not a computed value but an observed value — the runtime handler loaded on that node on the day of the accident, as confirmed with containerd config dump, was only runc.
Leave the judgment not to your hands but to the parser. Read the import list of the main configuration, expand the glob, and gather the names defined in the runtime tables on both sides. Only the last line is not a computation but an observed value from the day of the accident, and that value is written in the instructions.
A check script that looks at the loaded configuration
Create /root/gpuop/bin/check-runtime.sh. It must call containerd config dump without an absolute path, look for containerd.runtimes.nvidia in that output, and end with 0 if found and 1 if not. It must not read the /etc/containerd/config.toml path — the grader flags that as a wrong answer.
The basis of judgment is not the file but the configuration held by the running daemon. Call the command without an absolute path — the grader sets up a fake daemon at the front and actually runs your script twice. It must end with a nonzero value when it finds nothing for the check to do its job.
The RuntimeClass, the GPU Pod, and Pending
In /root/gpuop/k8s/runtimeclass.yaml, write a RuntimeClass whose name and handler are both nvidia, and in /root/gpuop/k8s/gpu-pod.yaml, write a Pod with the namespace gpu-lab, the name cuda-probe, runtimeClassName: nvidia, and nvidia.com/gpu: 1 in the container resource limits. Create the namespace first and apply both. The Pod does not come up — save the reason and message of the PodScheduled condition as two lines in /root/gpuop/out/unscheduled.txt.
A RuntimeClass is cluster-scoped so it has no namespace, and the name and the handler are different fields. It is normal for the Pod not to come up — leave the reason written in the scheduling condition in a file. It takes a few seconds for the condition to be filled in, so instead of sleeping for a fixed time, wait while checking the condition.
What the device plugin does, through the node status
With kubectl patch node lab-node-0 --subresource=status, put nvidia.com/gpu as "4" in both capacity and allocatable. When the Pod becomes Running, save the names and advertised amounts of the three nodes as a two-column table in /root/gpuop/out/node-gpu.txt.
Extended resources are under the status of the node object, and status is a separate subresource, so a plain patch does not take effect. You must fill in both capacity and allocatable, and write the value as a string. Once a slot appears, the scheduler picks up the waiting Pod again.
Read the time-slicing configuration and compute slots and memory
In /root/gpuop/k8s/time-slicing.yaml, write a ConfigMap with the namespace gpu-operator and the name time-slicing-config. data.any is a YAML string, and inside it there must be sharing.timeSlicing.failRequestsGreaterThanOne: true and the first item of the resources array must be name: nvidia.com/gpu with replicas: 5. After applying it, raise the advertised amount of lab-node-1 to "20", and write seven lines in /root/gpuop/out/slots.txt — PHYSICAL_GPUS REPLICAS ADVERTISED VRAM_PER_GPU_GIB VRAM_GUARANTEED_PER_SLOT_GIB VRAM_SHARED_PER_GPU_GIB ISOLATION. The equipment is assumed to be 4 A100s of 40GiB.
The data values of a ConfigMap must be strings, so if you write the configuration as a mapping, the apply is rejected. Put it in as a block scalar. The trap in the computation is the memory line — slots becoming five times as many does not mean the memory is divided into five.
Stall the rollout and report the split state
In /root/gpuop/k8s/toolkit-ds.yaml, write and apply a DaemonSet with the namespace gpu-operator, the name nvidia-container-toolkit, updateStrategy.rollingUpdate.maxUnavailable: 1, the image nvcr.io/nvidia/k8s/container-toolkit:v1.16.2, and a memory request of 128Mi. When all three Pods are Ready, raise the image to v1.17.0 and change the memory request to 64Gi at the same time. When the rollout stalls, write six lines in /root/gpuop/out/rollout.txt — DESIRED UPDATED READY OLD_IMAGE_PODS MAX_UNAVAILABLE, and BLOCKER=insufficient-memory.
First wait until all three nodes are ready, and only then push the new spec. Otherwise you cannot tell what stalled. The node capacity is 32Gi, so if you request more memory than that, no node can accept that Pod. Read the numbers in the report straight from the DaemonSet status and write them.