Operator and operands — switching six DaemonSets on and off per node
In one line
The dozen or so Pods the GPU Operator brings up are not one program but several DaemonSets that the Operator reads from the ClusterPolicy and unfolds, and because each DaemonSet uses the nvidia.com/gpu.deploy.<이름> label (the placeholder is the operand name) as its nodeSelector, it can be turned on and off per node.
Why this was needed
To use a GPU in Kubernetes, a node needs several things in place. A driver that loads the kernel module, a container toolkit that registers the GPU runtime in the container runtime, a device plugin that tells kubelet about the nvidia.com/gpu resource, GPU Feature Discovery that turns hardware facts into labels, a DCGM Exporter that emits metrics, a validator that checks the installation actually works, and, if you use MIG, even a MIG manager.
If you do this by hand, you have to install, upgrade, and roll back in order on every node. The Operator pattern adds to this "write the desired state in one object and a controller makes it so". That object is the ClusterPolicy, and the things the controller reads it and creates are the operands.
If you sort out the terminology once, the rest gets easy. The Operator is a single controller Pod that watches the ClusterPolicy. The operands are the workloads the Operator creates and manages, that is, the DaemonSets listed above. When an accident happens, "the Operator died" and "the operands did not come up" are entirely different problems, and the places to look are different too.
How it works
The operands are almost all DaemonSets. A DaemonSet is not "one on every node" but "one on each node that matches a condition", and that condition is the nodeSelector. The GPU Operator hangs its own labels on it.
| Operand | The label that turns it on |
|---|---|
| Driver | nvidia.com/gpu.deploy.driver |
| Container toolkit | nvidia.com/gpu.deploy.container-toolkit |
| Device plugin | nvidia.com/gpu.deploy.device-plugin |
| GPU Feature Discovery | nvidia.com/gpu.deploy.gpu-feature-discovery |
| DCGM Exporter | nvidia.com/gpu.deploy.dcgm-exporter |
| Validator | nvidia.com/gpu.deploy.operator-validator |
There are three practical knobs that come out of this structure.
One, take a whole node out of the operands. The official documentation describes a way to keep operands from being placed on that node with the label nvidia.com/gpu.deploy.operands=false.
kubectl label nodes $NODE nvidia.com/gpu.deploy.operands=false
Two, take out just one operand. To keep only the driver from coming up, it is nvidia.com/gpu.deploy.driver=false. Setting a label to false and deleting it outright read differently to people but are the same to a nodeSelector — if the value is not exactly "true", the condition does not match. However, writing false leaves "it was deliberately excluded", so in operations you use this one.
Three, turn off operands across the whole cluster. This is not a label but on the ClusterPolicy (or Helm values) side. In an environment where the driver is already installed on the host, it is driver.enabled=false, and if the GPU runtime is already registered, it is toolkit.enabled=false.
If you take a label down, the DaemonSet controller deletes that node's Pod. This behavior, Pods vanishing though you never edited the DaemonSet, is surprising at first, but once you know the controller is continually matching the condition against reality, it is natural. Conversely, if you put a label up, a Pod appears within seconds.
There is an order
The operands presuppose one another. Roughly, it is this chain.
드라이버 → 커널 모듈이 올라가고 /dev 에 장치가 생긴다
컨테이너 툴킷 → 런타임에 nvidia 런타임이 등록된다
장치 플러그인 → kubelet 에 nvidia.com/gpu 자원을 광고한다
GFD → 모델·장수·드라이버 판을 라벨로 붙인다
검증기 → 앞의 것들이 실제로 되는지 확인한다
So when you attach labels by hand, combinations that break the chain get created. The nastiest is a node where only the device plugin runs without the toolkit. The device plugin advertises the resource, the scheduler sees that number and sends Pods, and the Pods are scheduled successfully. But the runtime does not attach the GPU to the container, so only the workload cannot grab the device. It is a failure where no error appears anywhere yet only the result is wrong, so whether you have a check that finds such combinations decides your response time.
What it looks like in the field
First, "it's a GPU node but the Pod doesn't come up". The order of checking is the label → the DaemonSet's desiredNumberScheduled → that node's Pod. The three spots each say something different. If the label is missing, nothing happened; if the label is right but desiredNumberScheduled is low, the selector or a taint is off; and if the number is right but there is no Pod, it could not come up on that node.
Second, you are fooled if you look only at the number. desiredNumberScheduled is the number of nodes the DaemonSet controller judged should have a Pod. A node with a taint that is not tolerated drops out of the candidates altogether, so it is not even counted in this number. So when judging "are all the operands that should be there present", you must look not at the number but at the Pods actually Running on that node.
Third, the habit of isolating just one node. When a node with one odd GPU turns up, instead of cordoning the whole node, there is the option of taking down only its operands. With a single label line you stop that node's operands, and since the resource is no longer advertised, workloads naturally go to other nodes. It is a way to detach one machine without shaking the cluster.
Fourth, nodes where the driver is already installed. On-premises, many teams bake the driver into the image. If the driver operand comes up on such a node, it tries to load the kernel module twice, which does no good. If the whole cluster is like that, driver.enabled=false; if only some nodes are, nvidia.com/gpu.deploy.driver=false on those nodes. The official documentation states clearly that the GPU Operator does not manage the lifecycle of a driver preinstalled on the host — which means driver upgrades on those nodes are your job.
Fifth, the honest limit of this environment. The lab environment has neither a GPU nor a GPU Operator, so you cannot apply a ClusterPolicy. So you write by hand the DaemonSets the Operator would have unfolded. Instead, the DaemonSet controller and the scheduler are real, so Pods appearing and disappearing according to labels and nodes with a taint dropping out of the candidates are not imitation but the real thing. Whether the driver was really installed cannot be seen, and so is not looked at.
Reference documents
- Getting started with the GPU Operator (operand placement labels and deployment scenarios): https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/getting-started.html
- GPU Operator installation and chart values (driver.enabled and toolkit.enabled): https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/install-gpu-operator.html
- GPU Operator overview: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/overview.html
- DaemonSet: https://kubernetes.io/docs/concepts/workloads/controllers/daemonset/
- Assigning Pods to nodes (nodeSelector): https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/
What you will do in the next lab
You place, by label, the five DaemonSets the Operator would have unfolded. You see that with two GPU nodes the driver comes up on only one node, confirm that Pods vanish when you take down one label, and deliberately create a node where only the device plugin runs without the toolkit by attaching only labels by hand. And you build a tool that judges "are all the operands that should be there present" and "is there a node whose order is off", and have it checked that it does not lie even when one more node joins.