TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

With no label, nothing happens — and nothing complains

Continue in TT Lab

Goal

You set up the three sets of labels that make a GPU node (NFD, GFD, and the Operator) by origin, learn the label value convention by being rejected by the API server itself, decide a Pod's placement with In, NotIn, Exists, Gt, and a list of terms, and build by hand a synchronizer that restores discovery results and a label convention checker.

Why it matters

The GPU Operator does not look directly at the cards plugged into a node. It looks at labels. nfd-worker finds the PCI vendor ID 0x10de and attaches feature.node.kubernetes.io/pci-10de.present=true, the Operator brings up operands only on nodes with that label, and once gpu-feature-discovery attaches more labels for the model, count, memory, and driver version, from then on people and the scheduler choose placement by those labels. The frightening thing about this chain is that it raises no error when broken. Without the label, operands do not come up; since they do not come up, the resource is not advertised; and with no advertisement, the Pod is just Pending. Nowhere does it say "because the label is missing". So the first reflex of someone handling a GPU cluster becomes reading the labels first with kubectl get node -o json, and the second reflex becomes carrying a small tool that checks the label convention. This lab makes both.

Steps

  1. Work in /root/gpunfd (export KUBECONFIG=/root/.kube/config). Attach labels to the three nodes to set up a "cluster where discovery has finished". Attach to lab-node-0 feature.node.kubernetes.io/pci-10de.present=true, feature.node.kubernetes.io/kernel-version.major=5, nvidia.com/gpu.present=true, nvidia.com/gpu.product=NVIDIA-A100-SXM4-40GB, nvidia.com/gpu.count=4, nvidia.com/gpu.memory=40537, and nvidia.com/cuda.driver.major=550. Attach to lab-node-1, with the same keys, pci-10de.present=true, kernel-version.major=5, gpu.present=true, gpu.product=NVIDIA-A10, gpu.count=2, gpu.memory=22731, and cuda.driver.major=535. To lab-node-2, attach only feature.node.kubernetes.io/kernel-version.major=5, and do not attach pci-10de.present or any label starting with nvidia.com/ (it is a node without a GPU). And write three lines in /root/gpunfd/out/sources.txt — which component attaches which label. Write one line each of feature.node.kubernetes.io/pci-10de.present=nfd-worker, nvidia.com/gpu.product=gpu-feature-discovery, and nvidia.com/gpu.deploy.driver=gpu-operator.
  2. Try putting the device name NVIDIA A100-SXM4-40GB in as a label value as it is — run kubectl patch node lab-node-1 -p '{"metadata":{"labels":{"nvidia.com/gpu.product":"NVIDIA A100-SXM4-40GB"}}}' and save its output, including standard error, in /root/gpunfd/out/reject.txt (being rejected is normal, and the label does not change). And create /root/gpunfd/sanitize.sh — it tidies the string it receives as an argument so it can be used as a label value and prints it as one line. There are three rules. (1) Replace every character that is not alphanumeric or - _ . with -, (2) if it exceeds 63 characters, cut it to 63, (3) if the two ends are not alphanumeric, strip them until an alphanumeric appears. After creating it, check that bash /root/gpunfd/sanitize.sh "NVIDIA A100-SXM4-40GB" gives NVIDIA-A100-SXM4-40GB.
  3. Create the namespace nfd-lab and write the Pod a100-job in /root/gpunfd/k8s/a100-job.yaml. The container name is trainer and the image is nvcr.io/nvidia/pytorch:24.07-py3. Use a single spec.affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution to select only nodes where nvidia.com/gpu.product is NVIDIA-A100-SXM4-40GB. Do not use nodeSelector or nodeName. After applying it, check which node the Pod went to.
  4. Write the Pod any-gpu-job in /root/gpunfd/k8s/any-gpu-job.yaml (container trainer, same image). The condition is two matchExpressions inside one term — nodes where nvidia.com/gpu.present is Exists and, at the same time, nvidia.com/gpu.product is not (NotIn) NVIDIA-A10. After applying it, check that the Pod goes to lab-node-0. Note that you do not write values for Exists.
  5. Make two more Pods. gt-job in /root/gpunfd/k8s/gt-job.yaml has one condition in one term — nodes where nvidia.com/gpu.count is Gt 3. or-job in /root/gpunfd/k8s/or-job.yaml has two terms — the first is gpu.count Gt 3, and the second is gpu.product In [NVIDIA-A10]. Both have the container trainer and the same image. After applying them, confirm that gt-job can go only to lab-node-0 and or-job goes to one of the two GPU nodes, and write two lines in /root/gpunfd/out/placement.txt in the form <파드이름> <노드이름> (the placeholders are the Pod name and the node name).
  6. Under /root/gpunfd/features/, create one file per node, lab-node-0.env and lab-node-1.env. Each line is <라벨키>=<값> (the placeholders are the label key and the value) and holds, as they are, the nvidia.com/ labels you attached in step 1 (they are the facts the worker reported). And create /root/gpunfd/nfd-sync.sh — it cross-checks each file's lines against the actual node labels, and where they differ restores the file's value, printing for each restoration one line <노드> <키> <값> (the placeholders are the node, the key, and the value). If all are equal, it prints nothing and ends with 0. After creating it, spoil one value with kubectl label node lab-node-1 nvidia.com/gpu.count=9 --overwrite, run bash /root/gpunfd/nfd-sync.sh, and save its output to /root/gpunfd/out/sync.txt. When done, the label must be back to its original value.
  7. Attach feature.node.kubernetes.io/custom-gpu-training=true to lab-node-1 (suppose an NFD custom rule attached it). Write the Pod train-a in /root/gpunfd/k8s/train-a.yaml (container trainer, same image), make it require, with a required nodeAffinity, nodes where that label is true, apply it, and confirm that it comes up on lab-node-1. Then delete that label from lab-node-1, create a manifest with the same contents and only the name changed to train-b as /root/gpunfd/k8s/train-b.yaml, and apply it. Finally, write exactly two lines in /root/gpunfd/out/ignored.txt — train-a=Running:lab-node-1 and train-b=Pending.
  8. Create /root/gpunfd/label-audit.sh. It goes through all nodes in the cluster and, for nodes where nvidia.com/gpu.present is true, looks at four things — (1) are nvidia.com/gpu.product, nvidia.com/gpu.count, and nvidia.com/gpu.memory all present, (2) are gpu.count and gpu.memory made up of digits only, (3) is gpu.product at most 63 characters, (4) does gpu.product start with an alphanumeric and consist only of alphanumerics and - _ .. For each violation it prints one line <노드> <라벨키> <이유> (the placeholders are the node, the label key, and the reason), and if there is even one, it ends with exit code 1. If there is no problem, it prints only the single line OK and ends with 0. For the reason, write one of missing, not-a-number, too-long, and bad-format. After creating it, run it on the current cluster and save its output to /root/gpunfd/out/audit.txt. Do not write the node list inside the script — the grader creates one more node and then calls it.

Notes

Set up the discovery results — the three sets of labels have different origins

Work in /root/gpunfd (export KUBECONFIG=/root/.kube/config). Attach labels to the three nodes to set up a "cluster where discovery has finished". Attach to lab-node-0 feature.node.kubernetes.io/pci-10de.present=true, feature.node.kubernetes.io/kernel-version.major=5, nvidia.com/gpu.present=true, nvidia.com/gpu.product=NVIDIA-A100-SXM4-40GB, nvidia.com/gpu.count=4, nvidia.com/gpu.memory=40537, and nvidia.com/cuda.driver.major=550. Attach to lab-node-1, with the same keys, pci-10de.present=true, kernel-version.major=5, gpu.present=true, gpu.product=NVIDIA-A10, gpu.count=2, gpu.memory=22731, and cuda.driver.major=535. To lab-node-2, attach only feature.node.kubernetes.io/kernel-version.major=5, and do not attach pci-10de.present or any label starting with nvidia.com/ (it is a node without a GPU). And write three lines in /root/gpunfd/out/sources.txt — which component attaches which label. Write one line each of feature.node.kubernetes.io/pci-10de.present=nfd-worker, nvidia.com/gpu.product=gpu-feature-discovery, and nvidia.com/gpu.deploy.driver=gpu-operator.

You can attach several at once with kubectl label node <이름> <키>=<값> --overwrite (the placeholders are the name, the key, and the value). The prefix of a label key is its origin — feature.node.kubernetes.io/ is Node Feature Discovery's, discovery labels starting with nvidia.com/gpu. are GPU Feature Discovery's, and nvidia.com/gpu.deploy. is attached by the GPU Operator to gate its own operands. 0x10de is the PCI vendor ID assigned to NVIDIA.

A name the vendor gives cannot become a label as it is

Try putting the device name NVIDIA A100-SXM4-40GB in as a label value as it is — run kubectl patch node lab-node-1 -p '{"metadata":{"labels":{"nvidia.com/gpu.product":"NVIDIA A100-SXM4-40GB"}}}' and save its output, including standard error, in /root/gpunfd/out/reject.txt (being rejected is normal, and the label does not change). And create /root/gpunfd/sanitize.sh — it tidies the string it receives as an argument so it can be used as a label value and prints it as one line. There are three rules. (1) Replace every character that is not alphanumeric or - _ . with -, (2) if it exceeds 63 characters, cut it to 63, (3) if the two ends are not alphanumeric, strip them until an alphanumeric appears. After creating it, check that bash /root/gpunfd/sanitize.sh "NVIDIA A100-SXM4-40GB" gives NVIDIA-A100-SXM4-40GB.

The label value convention is decided by Kubernetes — at most 63 characters, starting and ending with an alphanumeric, and only - _ . in the middle. So gpu-feature-discovery does not use the device name the driver gives as it is but tidies it before attaching. You can replace it all at once with sed 's/[^A-Za-z0-9_.-]/-/g', and the cleanup of the ends can be done in two passes such as sed -E 's/^[^A-Za-z0-9]+//'. You must cut first and clean the ends afterwards for the convention to hold even when the 63rd character is a hyphen.

Choose a slot by model name

Create the namespace nfd-lab and write the Pod a100-job in /root/gpunfd/k8s/a100-job.yaml. The container name is trainer and the image is nvcr.io/nvidia/pytorch:24.07-py3. Use a single spec.affinity.nodeAffinity.requiredDuringSchedulingIgnoredDuringExecution to select only nodes where nvidia.com/gpu.product is NVIDIA-A100-SXM4-40GB. Do not use nodeSelector or nodeName. After applying it, check which node the Pod went to.

nodeSelector only looks at whether key and value are exactly equal. nodeAffinity does the same job while allowing In, NotIn, Exists, DoesNotExist, Gt, and Lt, so you can write conditions such as "one of these models". requiredDuringScheduling is a condition at scheduling time, and the trailing IgnoredDuringExecution means it is not applied again to Pods that are already up. You see where the Pod went with -o wide or .spec.nodeName.

The conditions within one term must all be true

Write the Pod any-gpu-job in /root/gpunfd/k8s/any-gpu-job.yaml (container trainer, same image). The condition is two matchExpressions inside one term — nodes where nvidia.com/gpu.present is Exists and, at the same time, nvidia.com/gpu.product is not (NotIn) NVIDIA-A10. After applying it, check that the Pod goes to lab-node-0. Note that you do not write values for Exists.

The matchExpressions within one nodeSelectorTerms item (term) must all be satisfied (AND). So a condition such as "among nodes that have a GPU, everything except this model" is written as one term. Exists and DoesNotExist do not look at values, so if you write values, the API server rejects it. The trap is that NotIn also lets through nodes that do not have that key at all — so you narrow to GPU nodes by also putting Exists.

A list of terms is OR, and numeric labels can be compared by size

Make two more Pods. gt-job in /root/gpunfd/k8s/gt-job.yaml has one condition in one term — nodes where nvidia.com/gpu.count is Gt 3. or-job in /root/gpunfd/k8s/or-job.yaml has two terms — the first is gpu.count Gt 3, and the second is gpu.product In [NVIDIA-A10]. Both have the container trainer and the same image. After applying them, confirm that gt-job can go only to lab-node-0 and or-job goes to one of the two GPU nodes, and write two lines in /root/gpunfd/out/placement.txt in the form <파드이름> <노드이름> (the placeholders are the Pod name and the node name).

nodeSelectorTerms is a list and among its items it is OR — if even one matches, that node becomes a candidate. Gt and Lt can be used only when the value reads as an integer, and you write just one value in values. Label values are strings, yet only these two operators interpret them as integers, which is a part often forgotten. When OR gives two candidates, the scheduler's score decides which way it goes — so "either one" is normal.

Why does a label you changed by hand come back

Under /root/gpunfd/features/, create one file per node, lab-node-0.env and lab-node-1.env. Each line is <라벨키>=<값> (the placeholders are the label key and the value) and holds, as they are, the nvidia.com/ labels you attached in step 1 (they are the facts the worker reported). And create /root/gpunfd/nfd-sync.sh — it cross-checks each file's lines against the actual node labels, and where they differ restores the file's value, printing for each restoration one line <노드> <키> <값> (the placeholders are the node, the key, and the value). If all are equal, it prints nothing and ends with 0. After creating it, spoil one value with kubectl label node lab-node-1 nvidia.com/gpu.count=9 --overwrite, run bash /root/gpunfd/nfd-sync.sh, and save its output to /root/gpunfd/out/sync.txt. When done, the label must be back to its original value.

NFD is the owner of the labels it attached, so the worker reports periodically and the master aligns the node with that report. So a value a person edited by hand quietly disappears at the next cycle — if you do not know the cause, you chase the ghost "the label keeps coming back". When reading labels, // in jq replaces even a false value with the default, so it is safe to branch with if . == null. To restore, use kubectl label --overwrite. It is idempotent only if there is no output on the second run.

Even if you delete the label, Pods already up are not evicted

Attach feature.node.kubernetes.io/custom-gpu-training=true to lab-node-1 (suppose an NFD custom rule attached it). Write the Pod train-a in /root/gpunfd/k8s/train-a.yaml (container trainer, same image), make it require, with a required nodeAffinity, nodes where that label is true, apply it, and confirm that it comes up on lab-node-1. Then delete that label from lab-node-1, create a manifest with the same contents and only the name changed to train-b as /root/gpunfd/k8s/train-b.yaml, and apply it. Finally, write exactly two lines in /root/gpunfd/out/ignored.txt — train-a=Running:lab-node-1 and train-b=Pending.

The name of the condition tells you the answer. requiredDuringSchedulingIgnoredDuringExecution demands it only at scheduling time, and during execution it does nothing even if the condition breaks. So even if you delete a label by mistake, existing Pods are fine, and from the next Pods to come up, they quietly become Pending — which is why the accident comes to light hours later. Deleting a label is kubectl label node <이름> <키>- (the placeholders are the name and the key). The reason for Pending is written in the PodScheduled message of .status.conditions.

Build a tool that checks whether the convention is kept

Create /root/gpunfd/label-audit.sh. It goes through all nodes in the cluster and, for nodes where nvidia.com/gpu.present is true, looks at four things — (1) are nvidia.com/gpu.product, nvidia.com/gpu.count, and nvidia.com/gpu.memory all present, (2) are gpu.count and gpu.memory made up of digits only, (3) is gpu.product at most 63 characters, (4) does gpu.product start with an alphanumeric and consist only of alphanumerics and - _ .. For each violation it prints one line <노드> <라벨키> <이유> (the placeholders are the node, the label key, and the reason), and if there is even one, it ends with exit code 1. If there is no problem, it prints only the single line OK and ends with 0. For the reason, write one of missing, not-a-number, too-long, and bad-format. After creating it, run it on the current cluster and save its output to /root/gpunfd/out/audit.txt. Do not write the node list inside the script — the grader creates one more node and then calls it.

Fetch the node list as you go with kubectl get nodes -o jsonpath='{range .items[*]}{.metadata.name}{"\n"}{end}'. If you fetch kubectl get node <이름> -o json (the placeholder is the name) just once per node and extract from it several times with jq, it fits comfortably within the 60-second budget. "The label is missing" and "the label value is an empty string" are different events, so you must not lump them together with // of jq. In the shell, count string length with ${#변수} (the placeholder is the variable). Pin the exit code with exit, not with the last echo.