TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

Why the GPU Operator Exists

Continue in TT Lab

In one line

The GPU Operator is not something that makes GPUs faster. It is moving the half-dozen installations you used to do by hand on every node into a single cluster object, and the price of that convenience is that the Operator now edits the node's container runtime configuration on your behalf.

Why this was needed

The procedure for making one GPU node "ready to use" used to go like this.

  1. Install matching kernel headers, then build and load the driver. If that fails, look at the kernel version first.
  2. Install nvidia-container-toolkit.
  3. Register the nvidia runtime handler in the containerd configuration.
  4. Restart containerd.
  5. Bring up nvidia-device-plugin so that the GPUs are advertised as a resource.
  6. Bring up dcgm-exporter to scrape utilization.

One node takes 30 minutes, and seven nodes take a day. The problem is not the time but drift. A node built in March has driver 535, and a node added in August has 550. When automatic kernel updates run, the driver breaks on only that node. A procedure that people repeat inevitably diverges, and a diverged cluster produces the question "why does it only fail on that node" every week.

The GPU Operator turns this whole procedure into a declaration of the desired state. If you write a single custom resource called ClusterPolicy, the Operator sets up the necessary DaemonSets, and when a new node joins, it goes through the same procedure by itself. That the result is the same even if you rebuild a node — that is the only product this Operator sells.

How it works

Under one ClusterPolicy, several DaemonSets are set up. Each takes one step of the procedure above.

DaemonSet What it does Symptom without it
nvidia-driver-daemonset Builds and loads the driver inside a container There is no nvidia-smi at all
nvidia-container-toolkit-daemonset Registers the runtime handler in the node's containerd configuration Pod creation fails with a RuntimeHandler error
nvidia-device-plugin-daemonset Advertises nvidia.com/gpu in the node status Pods stay Pending forever
gpu-feature-discovery Turns the GPU model and memory into node labels You cannot use scheduling conditions through labels
nvidia-dcgm-exporter Exposes utilization, temperature, and ECC metrics You do not know who is holding the GPU

There is an ordering dependency here. If the toolkit comes up before the driver is ready, it is meaningless, and if the device plugin advertises GPUs before the toolkit registers the runtime, the scheduler sends Pods but the node cannot create them. So the Operator gates the steps with node labels (nvidia.com/gpu.deploy.*). It is a mechanism so that people do not have to remember the order.

And there is one fact you must hold on to here. DaemonSet number 2 edits the node's /etc/containerd/ directly. It does not create an object inside the cluster; it writes to the host filesystem. The entire other half of this course is a story that comes out of that one sentence.

What it looks like in the field

First, the driver is usually installed on the node beforehand. Many organizations set driver.enabled=false and bake the driver into the OS image. The approach of building the driver inside a container creates, every time the kernel goes up, a stretch of several minutes in which the node runs without a GPU, and in an air-gapped network you are blocked already at downloading the header packages. Then what the GPU Operator does shrinks in effect to three things: "toolkit + device plugin + metrics".

Second, the version matrix is the real constraint. The driver, CUDA, toolkit, Operator, and Kubernetes each have a supported range. Upgrading the Operator means moving those five at once, so if you put it in the same window as other upgrades, you cannot separate the causes.

Third, from an audit point of view this Operator is privileged software. It writes to the node's filesystem, sees the host PID namespace, and loads kernel modules. The intuition "it is a Pod running on Kubernetes, so it is isolated" does not hold here. If you do not leave this fact in the documentation when approving the installation, it will certainly become a problem later.

What to look at in the next reading

In the next article, you see how the DaemonSets under ClusterPolicy connect to a RuntimeClass and the nvidia.com/gpu resource. These two strands are entirely different gates, and if you confuse them, you fall straight into the state of "the Pod is Pending and I cannot find the cause".