TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

What Is Real and What Is Imitation

Continue in TT Lab

In one line

This lab Pod has no real GPU, no real containerd, and no GPU Operator. Instead it has a TOML parser and a real Kubernetes control plane, and almost everything that caught people in this accident can be reproduced with those two.

Why do this in a fake environment

If you count, one by one, what you actually need to learn from a GPU accident, it goes like this.

nvidia-smi is not on this list. What you can learn only by actually touching a GPU device is driver builds and real kernel execution performance, and neither of those was the cause of this accident. The causes were all in configuration and objects.

What is truly validated and what is imitation

What you do in the lab Its status in this environment
Writing containerd configuration as TOML and parsing it Real. A real parser reads it
Judging table collisions between the main configuration and a drop-in Real. The computation is correct as it is
Verifying the behavior of the check script Real. A fake containerd is set up and it is actually run twice
RuntimeClass registration Real. A real API server stores it
GPU resource advertisement and scheduling Real. A real scheduler judges it
The stall of a DaemonSet rolling update Real. A real controller stops
A GPU actually being plugged into a node Imitation. Only a number is written in the node status
Using a GPU inside a container None. Containers do not run
Restarting containerd so that a handler gets registered None. It is dealt with only through concepts and the check script

Let us make the last row clear in particular. The fact that SIGHUP is not enough and a restart is needed cannot be demonstrated in this Pod. So the lab has you create, instead, "a check script that tells the truth regardless of whether there was a restart". The spot where that knowledge comes out in your hands in the field is exactly that script.

What is used in the field as it is

Three of the deliverables you make in the lab can be taken to the company as they are.

What changes on a real GPU node

If you know how the field differs in the parts this lab left as imitation, you will not be flustered when you stand in front of the real thing later.

The driver and the kernel move together. A GPU driver is a kernel module, so when the kernel version changes it must be built again. So if the node is configured to update the kernel automatically, the node comes up Ready after a reboot with the GPU gone. Pods are scheduled and containers come up, but only the device is missing, so the symptom appears only as an application error, "cannot find the GPU". The standard defense is to attach the driver version to a node label and block scheduling when it disappears.

Resource advertisement goes through several layers. The device plugin registers with kubelet, kubelet posts to the node status, and the scheduler looks at it. If it is cut at any layer, the result is the same: "the Pod is Pending". So the investigation goes from the top down. Is the resource in the node status → if not, the kubelet log → the device plugin Pod state → did that Pod mount the socket directory correctly. If you do not fix this order, you end up digging from a different place each time.

There are several ways to share a GPU. Time-slicing rotates turns, so memory is not isolated, and if one Pod uses all the memory, the rest die with it. MIG divides in hardware, so isolation is certain, but the combinations you can divide into are fixed and you must empty the node when you change them. MIG if you need isolation, time-slicing if utilization is the goal is the practical baseline, and you must be able to explain this choice to users.

To sum up, what you get a feel for in this lab is the order of judgment and the way you leave your evidence. Those two apply as they are regardless of whether the GPU is real or an imitation, and what you must newly learn in front of the real thing narrows to about the three paragraphs above.

What you will do in the next lab

You walk through the accident again from the start in eight steps. The first four steps deal with the node's configuration files, and the last four steps deal with cluster objects. The last two steps gather what you made earlier and finish with computation and reporting.