TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

Node Feature Discovery — how a discovered fact becomes a label

Continue in TT Lab

In one line

The GPU Operator does not look directly at the cards plugged into a node; it looks at node labels, and those labels are attached in shares by Node Feature Discovery (NFD), GPU Feature Discovery (GFD), and the Operator itself. So if a label is missing, nothing happens at all, with neither error nor warning.

Why this was needed

Software that handles GPU nodes must know at least three things. Does this node have an NVIDIA card, how many and which model, and which version is the driver. The way to find this out used to be to go into the node, type lspci, and run nvidia-smi. With three nodes it is doable, and with three hundred it is impossible.

Kubernetes' answer to this problem was "let's put hardware facts up as labels on the node object". Once it becomes a label, everything after is solved with Kubernetes' everyday tools. The scheduler's nodeSelector and nodeAffinity read that label, a DaemonSet's nodeSelector decides placement with that label, and people see it at a glance with kubectl get node -L. The problem of querying hardware turned into the problem of querying labels — that is the whole of this design.

NFD is the Kubernetes SIG project that does this job. Its structure is split in two. The nfd-worker, running as a DaemonSet on each node, gathers and reports facts in strands such as CPU, kernel, PCI, USB, memory, storage, and network, and the nfd-master receives those reports and writes labels on the node object. The reason the worker goes through the master instead of editing the node directly is to concentrate the permission to edit node objects in one place.

How it works

Labels reveal their origin by prefix. Remembering just this one thing changes how fast you read the screen.

Prefix Who attaches it Examples
feature.node.kubernetes.io/ nfd-master, which received the report from nfd-worker pci-10de.present, kernel-version.full, system-os_release.ID
nvidia.com/gpu.*, nvidia.com/cuda.* gpu-feature-discovery gpu.product, gpu.count, gpu.memory, cuda.driver.major
nvidia.com/gpu.deploy.* The GPU Operator gpu.deploy.driver, gpu.deploy.container-toolkit

The NVIDIA documentation says it identifies a GPU worker node by the presence of the label feature.node.kubernetes.io/pci-10de.present=true. 0x10de is the PCI vendor ID assigned to NVIDIA. NFD's PCI labels by default use device names of the form <class>_<vendor>, but the GPU Operator configures them to look at the vendor only.

The labels GFD attaches are more specific. If you transcribe the output example of the official documentation as it is, it looks like this.

{
  "nvidia.com/cuda.driver.major": "450",
  "nvidia.com/cuda.driver.minor": "80",
  "nvidia.com/cuda.driver.rev": "02",
  "nvidia.com/cuda.runtime.major": "11",
  "nvidia.com/cuda.runtime.minor": "0",
  "nvidia.com/gpu.compute.major": "8",
  "nvidia.com/gpu.count": "1",
  "nvidia.com/gpu.family": "ampere",
  "nvidia.com/gpu.memory": "40537",
  "nvidia.com/gpu.product": "A100-SXM4-40GB"
}

It is worth noticing that the driver version is split into three pieces, major, minor, and rev. The reason 450.80.02 was not put into a single label whole is comparison. The Gt and Lt operators of nodeAffinity read values as integers, so only if they are split into pieces can you write a condition such as "nodes with driver 550 or higher".

Label values have syntax constraints. At most 63 characters, starting and ending with an alphanumeric, and only -, _, and . can come in the middle. But the device name the driver reports contains a space, as in NVIDIA A100-SXM4-40GB. If you try to put it in as it is, the API server rejects it with Invalid value. So GFD tidies the name before attaching it, and the NVIDIA-A100-SXM4-40GB we see on the screen is the result of that normalization. If you do not know this, you spend a long time at the spot of "I set a condition with the model name written in the documentation but no node matches".

The side that reads labels has rules too. nodeSelector only looks at whether key and value are exactly equal, but nodeAffinity can use In, NotIn, Exists, DoesNotExist, Gt, and Lt. And there is one structure that people often get wrong.

And NotIn also lets through nodes that do not have that key at all. If you write "on GPU nodes except A10" with a single NotIn, even nodes with no GPU become candidates. That is why it became an idiom to narrow the scope first by also putting Exists.

The IgnoredDuringExecution at the end of the name is not just tacked on, either. requiredDuringSchedulingIgnoredDuringExecution demands the condition only at scheduling time, and does not evict Pods that are already up even if the condition breaks. So even if you delete a label by mistake, workloads that were running are fine, and from the next Pods to come up, they quietly become Pending. It is a structure in which the accident comes to light hours later.

Ownership of labels

The labels NFD attached are NFD's. The worker reports again periodically and the master aligns the node with that report, so a value a person edited by hand quietly reverts at the next cycle. The first time you meet it, you chase the ghost "the label keeps coming back to life", but once you know, it is obviously expected behavior — if a person could overwrite a discovered fact, there would be no reason to trust that label.

So facts decided by people are written under a different key. Things such as team ownership, workload tier, and maintenance window go in labels using your own company prefix, and discovery labels are only read. Conversely, if you want NFD to carry out a person's rule, you declare the rule with NodeFeatureRule and have NFD attach it — then that label also becomes NFD's and is managed consistently.

What it looks like in the field

First, nothing happens. The most common cause of the report "I installed the GPU Operator but not a single operand comes up" is the absence of the pci-10de.present label. Either you installed NFD separately and the label prefix setting changed, or a node joined newly and the nfd-worker DaemonSet could not come up on it. Either way, there is no error log. If the condition does not match, placement does not happen, and things that do not happen leave no log.

Second, the model name is subtly different. nvidia.com/gpu.product is a value that has gone through GFD's normalization, and if you turn on MIG, a suffix is attached, as in NVIDIA-H100-80GB-HBM3-MIG-1g.10gb. If you turn on time-slicing, yet another suffix is attached. So when you set a condition by model name, you must look at the label value in the cluster right now, not the name in the documentation.

Third, teams that carry a label checker are faster. It really happens that a node claims to be a GPU node while one of model, count, and memory is empty, or that a value that should be a number is not a number. If you catch this with a small script instead of by eye, the check is over in 30 seconds every time a node joins. What matters is that the script fetches the node list as it goes. A checker with node names baked in starts lying from the day you add capacity.

Fourth, the honest limit of this environment. The lab environment has no GPU, no NFD, and no GFD. So "discovery" is something you set up with labels. Instead, label value validation, nodeAffinity evaluation, the scheduler's placement decisions, and adding nodes are done by the real API server and real scheduler brought up by kwok — not an imitation but the real thing.

Reference documents

What you will do in the next lab

On the real API server brought up by kwok, you set up labels from three origins by hand. You put in a device name containing a space and get rejected by the API server, and then make a tool that tidies that name into a label value. With In, NotIn, Exists, Gt, and a list of terms, you have the scheduler itself confirm which node a Pod goes to, and you see with your own eyes that Pods that were running remain after the label is deleted and only new Pods become Pending. Finally, you make a synchronizer that restores the discovery results and a label convention checker that does not lie even when new nodes join.