What Is Real and What Is Imitation
In one line
This lab Pod has no real GPU, no real containerd, and no GPU Operator. Instead it has a TOML parser and a real Kubernetes control plane, and almost everything that caught people in this accident can be reproduced with those two.
Why do this in a fake environment
If you count, one by one, what you actually need to learn from a GPU accident, it goes like this.
- How to judge whether two configuration files collide over the same table → a parser is enough
- How to write a check script that looks at what is loaded rather than the file → a shell is enough
- Where a Pod gets blocked when there is no resource advertisement → the scheduler must be real
- How to compute the slot count and memory in a time-slicing configuration → arithmetic is enough
- What a rollout stalled on one node looks like → the DaemonSet controller must be real
nvidia-smi is not on this list. What you can learn only by actually touching a GPU device is driver builds and real kernel execution performance, and neither of those was the cause of this accident. The causes were all in configuration and objects.
What is truly validated and what is imitation
| What you do in the lab | Its status in this environment |
|---|---|
| Writing containerd configuration as TOML and parsing it | Real. A real parser reads it |
| Judging table collisions between the main configuration and a drop-in | Real. The computation is correct as it is |
| Verifying the behavior of the check script | Real. A fake containerd is set up and it is actually run twice |
| RuntimeClass registration | Real. A real API server stores it |
| GPU resource advertisement and scheduling | Real. A real scheduler judges it |
| The stall of a DaemonSet rolling update | Real. A real controller stops |
| A GPU actually being plugged into a node | Imitation. Only a number is written in the node status |
| Using a GPU inside a container | None. Containers do not run |
| Restarting containerd so that a handler gets registered | None. It is dealt with only through concepts and the check script |
Let us make the last row clear in particular. The fact that SIGHUP is not enough and a restart is needed cannot be demonstrated in this Pod. So the lab has you create, instead, "a check script that tells the truth regardless of whether there was a restart". The spot where that knowledge comes out in your hands in the field is exactly that script.
What is used in the field as it is
Three of the deliverables you make in the lab can be taken to the company as they are.
check-runtime.sh— hang it as it is on the post-boot check of a node or on a CI gate. The key is that the basis of judgment is the dump, not the file, and that property is independent of the environment.- The slot calculation table — a table to show users before turning on time-slicing. The two lines "20 slots, guaranteed memory 0" settle half the conversation.
- The rollout stall report — it leaves which column you looked at and what you judged. This is the skeleton of an incident report.
What changes on a real GPU node
If you know how the field differs in the parts this lab left as imitation, you will not be flustered when you stand in front of the real thing later.
The driver and the kernel move together. A GPU driver is a kernel module, so when the kernel version changes it must be built again. So if the node is configured to update the kernel automatically, the node comes up Ready after a reboot with the GPU gone. Pods are scheduled and containers come up, but only the device is missing, so the symptom appears only as an application error, "cannot find the GPU". The standard defense is to attach the driver version to a node label and block scheduling when it disappears.
Resource advertisement goes through several layers. The device plugin registers with kubelet, kubelet posts to the node status, and the scheduler looks at it. If it is cut at any layer, the result is the same: "the Pod is Pending". So the investigation goes from the top down. Is the resource in the node status → if not, the kubelet log → the device plugin Pod state → did that Pod mount the socket directory correctly. If you do not fix this order, you end up digging from a different place each time.
There are several ways to share a GPU. Time-slicing rotates turns, so memory is not isolated, and if one Pod uses all the memory, the rest die with it. MIG divides in hardware, so isolation is certain, but the combinations you can divide into are fixed and you must empty the node when you change them. MIG if you need isolation, time-slicing if utilization is the goal is the practical baseline, and you must be able to explain this choice to users.
To sum up, what you get a feel for in this lab is the order of judgment and the way you leave your evidence. Those two apply as they are regardless of whether the GPU is real or an imitation, and what you must newly learn in front of the real thing narrows to about the three paragraphs above.
What you will do in the next lab
You walk through the accident again from the start in eight steps. The first four steps deal with the node's configuration files, and the last four steps deal with cluster objects. The last two steps gather what you made earlier and finish with computation and reporting.