One label went down and the pod was gone
Goal
You place by node label the five DaemonSets the GPU Operator unfolds, take out just one operand from one node, and build a tool that judges "are all the operands that should be there present" and "is there a node whose order is off".
Why it matters
When you install the GPU Operator, more than ten Pods come up. Those Pods are not one program but several DaemonSets that the Operator reads from the ClusterPolicy and unfolds. The driver, container toolkit, device plugin, GPU Feature Discovery, DCGM exporter, validator, and MIG manager each have their own DaemonSet, and each DaemonSet looks at the nvidia.com/gpu.deploy.<이름> label (the placeholder is the operand name) through its nodeSelector. If you know this structure, what you can do in operations changes — taking just one node out of the operands, not bringing up only the driver on a node where the driver is already installed, and isolating one problem node are each a single label line. If you do not know it, it goes the other way. If one label is attached wrongly and a node appears where only the device plugin runs without the toolkit, the resource is advertised and the scheduler succeeds, but only the workload cannot grab the device. Nobody raises an error, so hours just go by.
Steps
- Work in
/root/gpuops(export KUBECONFIG=/root/.kube/config). Create the namespacegpu-operator. Attach to lab-node-0 (a GPU node, no driver)feature.node.kubernetes.io/pci-10de.present=trueandnvidia.com/gpu.present=true, together withnvidia.com/gpu.deploy.driver=true,nvidia.com/gpu.deploy.container-toolkit=true,nvidia.com/gpu.deploy.device-plugin=true,nvidia.com/gpu.deploy.gpu-feature-discovery=true, andnvidia.com/gpu.deploy.dcgm-exporter=true. To lab-node-1 (a GPU node where the driver is already installed on the host), attach the same labels but leave onlynvidia.com/gpu.deploy.driverasfalse. To lab-node-2 (no GPU), attach not a single label starting withnvidia.com/gpu.deploy.. And write five lines in/root/gpuops/out/roster.txt— in the form<오퍼랜드이름>=<그 오퍼랜드를 켜는 라벨 키>(the placeholders are the operand name and the label key that turns that operand on), the five ofdriver,container-toolkit,device-plugin,gpu-feature-discovery, anddcgm-exporter. - In
/root/gpuops/k8s/toolkit-ds.yaml, write the DaemonSetnvidia-container-toolkit-daemonset— namespacegpu-operator, the Pod label and selectorapp: nvidia-container-toolkit-daemonset, thenodeSelectornvidia.com/gpu.deploy.container-toolkit: "true", the container imagenvcr.io/nvidia/k8s/container-toolkit:v1.16.2, and a memory request of128Mi. After applying it, confirm thatdesiredNumberScheduledbecomes 2 and that a Pod comes up on each of lab-node-0 and lab-node-1. - In
/root/gpuops/k8s/driver-ds.yaml, write the DaemonSetnvidia-driver-daemonset— the same namespace, the Pod label and selectorapp: nvidia-driver-daemonset, the nodeSelectornvidia.com/gpu.deploy.driver: "true", the imagenvcr.io/nvidia/driver:550.90.07-ubuntu22.04, and a memory request of512Mi. After applying it, confirm that even though there are two GPU nodes, only this DaemonSet comes up on one node, and write two lines in/root/gpuops/out/driver.txt—DESIRED=<숫자>andSKIPPED=<빠진 노드 이름>(the placeholders are the number and the name of the skipped node). - Make two more DaemonSets.
nvidia-device-plugin-daemonsetin/root/gpuops/k8s/device-plugin-ds.yamlhas the nodeSelectornvidia.com/gpu.deploy.device-plugin: "true", the imagenvcr.io/nvidia/k8s-device-plugin:v0.16.2, and a memory request of128Mi.gpu-feature-discoveryin/root/gpuops/k8s/gfd-ds.yamlhas the nodeSelectornvidia.com/gpu.deploy.gpu-feature-discovery: "true", the same image, and the same memory request. For both, the Pod label and selector isapp: <데몬셋 이름>(the placeholder is the DaemonSet name). After applying them, confirm that thedesiredNumberScheduledof each of the two DaemonSets is 2. - In
/root/gpuops/k8s/dcgm-ds.yaml, write the DaemonSetnvidia-dcgm-exporter— the nodeSelectornvidia.com/gpu.deploy.dcgm-exporter: "true", the imagenvcr.io/nvidia/k8s/dcgm-exporter:3.3.7-3.5.0-ubuntu22.04, a memory request of128Mi, and the Pod label and selectorapp: nvidia-dcgm-exporter. After applying it and confirming that it comes up on both nodes, changenvidia.com/gpu.deploy.dcgm-exporterof lab-node-1 tofalseand confirm that the Pod disappears. Write two lines in/root/gpuops/out/optout.txt—BEFORE=<바꾸기 전 desiredNumberScheduled>andAFTER=<바꾼 뒤 값>(the placeholders are the desiredNumberScheduled before the change and the value after the change). - Create
/root/gpuops/operand-audit.sh. It goes through all nodes in the cluster, and for each operand whosenvidia.com/gpu.deploy.<이름>label (the placeholder is the operand name) is attached to that node astrue, it checks whether that DaemonSet's Pod is Running on that node. If not, it prints one line<노드> MISSING <오퍼랜드이름>(the placeholders are the node and the operand name), and if nothing is missing on that node, it prints one line<노드> OK(the placeholder is the node). If even one is missing, the exit code is 1, and otherwise 0. The pairing of operand names and DaemonSet names is the five from step 1. After creating it, run it on the current cluster and save the output to/root/gpuops/out/audit.txt. Do not write node names inside the script — the grader creates one more node and then calls it. - Suppose someone attached only
nvidia.com/gpu.deploy.device-plugin=trueto lab-node-2 by hand. Actually attach that label (do not attach otherdeploylabels) and confirm that the device plugin Pod comes up on that node. And create/root/gpuops/order-check.sh <노드이름>(the placeholder is the node name) — if the device plugin Pod is Running on that node but there is no container toolkit Pod, it gives the one linedevice-plugin-without-toolkitand exit code 1, and in all other cases the one lineokand 0. Run it on the three nodes in turn and write three lines in/root/gpuops/out/order.txtin the form<노드> <결과>(the placeholders are the node and the result). - Write three lines in
/root/gpuops/out/matrix.txt, one per node. The format is<노드> driver=<yes|no> toolkit=<yes|no> device-plugin=<yes|no> gfd=<yes|no> dcgm=<yes|no>(the placeholder is the node), andyesmeans that operand's DaemonSet Pod is Running on that node. The line order is lab-node-0, lab-node-1, lab-node-2. And on a fourth line, writeTOTAL_PODS=<gpu-operator 네임스페이스에서 Running 인 데몬셋 파드 총수>(the placeholder is the total number of DaemonSet Pods Running in the gpu-operator namespace). Fill in the numbers and yes/no by asking the API — if you write from the memory of earlier steps, they will not match.
Notes
- Start with
export KUBECONFIG=/root/.kube/config. There are three nodes, lab-node-0/1/2, all deliverables go under/root/gpuops, and objects go in the namespacegpu-operator. - This environment has neither a GPU nor a GPU Operator. So you write by hand the DaemonSets the Operator would have created. Instead, the DaemonSet controller and the scheduler are real, so Pods appearing and disappearing according to labels is the real thing.
- Pods are not actually run and are forged to Ready. Whether the driver was really installed cannot be seen and is not looked at — all this lab judges is "what came up on which node".
- Right after applying a DaemonSet or changing a label, the controller needs a few seconds to react. If a number looks like the old one, look again a moment later.
- Common mistake: leaving out the quotation marks in the nodeSelector value. Label values are strings, so write
"true". - Common mistake: judging with
desiredNumberScheduledalone. A node where the label is right but the Pod could not come up is not caught by that number. - GPU Operator: Getting Started · Installing the NVIDIA GPU Operator · DaemonSet
Put the labels that turn operands on and off differently on each node
Work in /root/gpuops (export KUBECONFIG=/root/.kube/config). Create the namespace gpu-operator. Attach to lab-node-0 (a GPU node, no driver) feature.node.kubernetes.io/pci-10de.present=true and nvidia.com/gpu.present=true, together with nvidia.com/gpu.deploy.driver=true, nvidia.com/gpu.deploy.container-toolkit=true, nvidia.com/gpu.deploy.device-plugin=true, nvidia.com/gpu.deploy.gpu-feature-discovery=true, and nvidia.com/gpu.deploy.dcgm-exporter=true. To lab-node-1 (a GPU node where the driver is already installed on the host), attach the same labels but leave only nvidia.com/gpu.deploy.driver as false. To lab-node-2 (no GPU), attach not a single label starting with nvidia.com/gpu.deploy.. And write five lines in /root/gpuops/out/roster.txt — in the form <오퍼랜드이름>=<그 오퍼랜드를 켜는 라벨 키> (the placeholders are the operand name and the label key that turns that operand on), the five of driver, container-toolkit, device-plugin, gpu-feature-discovery, and dcgm-exporter.
There is one Operator and there are several operands — the Operator that read the ClusterPolicy unfolds them into DaemonSets such as the driver, toolkit, device plugin, GFD, and DCGM exporter. The nodeSelector of each DaemonSet looks at the nvidia.com/gpu.deploy.<이름> label (the placeholder is the operand name), so taking just one operand out of one node is done with a single label line. The official documentation gives nvidia.com/gpu.deploy.driver=false as the way to not bring up the driver on specific nodes only. You can attach several at once with kubectl label node <이름> <키>=<값> --overwrite (the placeholders are the name, the key, and the value).
Place the first operand by label
In /root/gpuops/k8s/toolkit-ds.yaml, write the DaemonSet nvidia-container-toolkit-daemonset — namespace gpu-operator, the Pod label and selector app: nvidia-container-toolkit-daemonset, the nodeSelector nvidia.com/gpu.deploy.container-toolkit: "true", the container image nvcr.io/nvidia/k8s/container-toolkit:v1.16.2, and a memory request of 128Mi. After applying it, confirm that desiredNumberScheduled becomes 2 and that a Pod comes up on each of lab-node-0 and lab-node-1.
A DaemonSet is not "one per node" but "one per node that matches the condition". That condition is the nodeSelector, and the GPU Operator hangs its own labels on it to turn operands on and off per node. desiredNumberScheduled is "the number of nodes the DaemonSet controller judged this DaemonSet should be on" — as the number of nodes with matching labels grows or shrinks, this number follows. You see it at a glance with kubectl -n gpu-operator get ds -o wide.
The node where only the driver is left out
In /root/gpuops/k8s/driver-ds.yaml, write the DaemonSet nvidia-driver-daemonset — the same namespace, the Pod label and selector app: nvidia-driver-daemonset, the nodeSelector nvidia.com/gpu.deploy.driver: "true", the image nvcr.io/nvidia/driver:550.90.07-ubuntu22.04, and a memory request of 512Mi. After applying it, confirm that even though there are two GPU nodes, only this DaemonSet comes up on one node, and write two lines in /root/gpuops/out/driver.txt — DESIRED=<숫자> and SKIPPED=<빠진 노드 이름> (the placeholders are the number and the name of the skipped node).
There are two GPU nodes but the driver DaemonSet comes up on only one. That is because if the label value is not "true", the nodeSelector does not match — writing false and deleting the label outright read differently to people but are the same to the nodeSelector. This is how you express "the driver is already installed on this node". The driver is an operand that loads a kernel module, so it must not overlap with a driver already installed on the host. Do not write the values by looking at command output; take them with -o jsonpath and write them.
Bring up the device plugin and GFD together
Make two more DaemonSets. nvidia-device-plugin-daemonset in /root/gpuops/k8s/device-plugin-ds.yaml has the nodeSelector nvidia.com/gpu.deploy.device-plugin: "true", the image nvcr.io/nvidia/k8s-device-plugin:v0.16.2, and a memory request of 128Mi. gpu-feature-discovery in /root/gpuops/k8s/gfd-ds.yaml has the nodeSelector nvidia.com/gpu.deploy.gpu-feature-discovery: "true", the same image, and the same memory request. For both, the Pod label and selector is app: <데몬셋 이름> (the placeholder is the DaemonSet name). After applying them, confirm that the desiredNumberScheduled of each of the two DaemonSets is 2.
The device plugin is the operand that advertises the nvidia.com/gpu resource to kubelet, and GFD is the operand that turns that node's GPU facts into labels and attaches them. The two come from the same image but do different jobs, so there are separate DaemonSets, and because they are separate, there are separate labels too. Even on the node where the driver was left out, these two must come up — because the driver is already on the host. Note that there are operands whose names do not start with nvidia-.
When you take the label down, the Pod disappears
In /root/gpuops/k8s/dcgm-ds.yaml, write the DaemonSet nvidia-dcgm-exporter — the nodeSelector nvidia.com/gpu.deploy.dcgm-exporter: "true", the image nvcr.io/nvidia/k8s/dcgm-exporter:3.3.7-3.5.0-ubuntu22.04, a memory request of 128Mi, and the Pod label and selector app: nvidia-dcgm-exporter. After applying it and confirming that it comes up on both nodes, change nvidia.com/gpu.deploy.dcgm-exporter of lab-node-1 to false and confirm that the Pod disappears. Write two lines in /root/gpuops/out/optout.txt — BEFORE=<바꾸기 전 desiredNumberScheduled> and AFTER=<바꾼 뒤 값> (the placeholders are the desiredNumberScheduled before the change and the value after the change).
The DaemonSet controller keeps matching the nodeSelector against node labels. When a label leaves the condition, it deletes that node's Pod — which is why Pods vanish though you did not edit the DaemonSet. The knob you use in operations when you want to take just one node out of the operands is exactly this. The official documentation also has nvidia.com/gpu.deploy.operands=false, which takes all of a node's operands out at once. You must take the two numbers before and after changing the label — you cannot write them all at once afterwards.
A tool that judges whether all the operands that should be there are present
Create /root/gpuops/operand-audit.sh. It goes through all nodes in the cluster, and for each operand whose nvidia.com/gpu.deploy.<이름> label (the placeholder is the operand name) is attached to that node as true, it checks whether that DaemonSet's Pod is Running on that node. If not, it prints one line <노드> MISSING <오퍼랜드이름> (the placeholders are the node and the operand name), and if nothing is missing on that node, it prints one line <노드> OK (the placeholder is the node). If even one is missing, the exit code is 1, and otherwise 0. The pairing of operand names and DaemonSet names is the five from step 1. After creating it, run it on the current cluster and save the output to /root/gpuops/out/audit.txt. Do not write node names inside the script — the grader creates one more node and then calls it.
This judgment must be made not by "how many does the DaemonSet want" but by "is it actually up on this node". The two differ — nodes do arise where the label is right but the Pod cannot come up because of a taint or resources. If you call kubectl for every node, it gets slower every time nodes grow. If you fetch the node list and the Pod list once each and filter with jq, the API calls end at two. If you use // when reading labels with jq, a label whose value is false and a missing label cannot be told apart.
Find the combination that makes no sense
Suppose someone attached only nvidia.com/gpu.deploy.device-plugin=true to lab-node-2 by hand. Actually attach that label (do not attach other deploy labels) and confirm that the device plugin Pod comes up on that node. And create /root/gpuops/order-check.sh <노드이름> (the placeholder is the node name) — if the device plugin Pod is Running on that node but there is no container toolkit Pod, it gives the one line device-plugin-without-toolkit and exit code 1, and in all other cases the one line ok and 0. Run it on the three nodes in turn and write three lines in /root/gpuops/out/order.txt in the form <노드> <결과> (the placeholders are the node and the result).
The container toolkit must register the runtime for a GPU container to see the device. If only the device plugin is up, the resource is advertised but the Pod that received that resource cannot grab the device — the scheduler says it succeeded yet the workload fails, the hardest combination to find. So you must look not at "what is up" but at "what is up together with what". The script looks only at the node it receives as an argument, and judges by the number of Pods of the two DaemonSets. A script that always gives 1 will not do — the grader calls it with a healthy node too.
Extract the table of nodes and operands from the actual state
Write three lines in /root/gpuops/out/matrix.txt, one per node. The format is <노드> driver=<yes|no> toolkit=<yes|no> device-plugin=<yes|no> gfd=<yes|no> dcgm=<yes|no> (the placeholder is the node), and yes means that operand's DaemonSet Pod is Running on that node. The line order is lab-node-0, lab-node-1, lab-node-2. And on a fourth line, write TOTAL_PODS=<gpu-operator 네임스페이스에서 Running 인 데몬셋 파드 총수> (the placeholder is the total number of DaemonSet Pods Running in the gpu-operator namespace). Fill in the numbers and yes/no by asking the API — if you write from the memory of earlier steps, they will not match.
This table is precisely "what shape the GPU stack of this cluster is in right now". It is the first thing you make when an accident happens, and if you have the command that makes the table, it is over in 30 seconds. One kubectl -n gpu-operator get pods -o json contains everything you need — look at .spec.nodeName, .metadata.labels.app, and .status.phase together. On lab-node-2, because of the label you attached in step 7, only the device plugin should be up.