TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

Sandbox workloads — the label that picks the resource name, and the inventory that splits

Continue in TT Lab

Goal

You set up three nodes for different GPU workload uses, confirm with the real scheduler that the resource name changes per use and that as a result the inventory splits, and then build a checker that judges whether labels and advertisements agree.

Why it matters

Not all workloads that use GPUs are containers. Some must run inside virtual machines because of licenses or legacy, and to hand a card to such a VM, even the driver installed on the node must differ — the data center driver for containers, vfio-pci for passthrough, and the vGPU Manager for vGPU. The GPU Operator accepts that choice as a single node label line. And the result of that choice changes even the resource name the node advertises. If the names differ, they are different resources to the scheduler, so even when the queue on one side is long, the slots on the other side stay empty. This is where it comes from that splitting inventory by use is a hard decision to reverse.

Steps

  1. Create /root/gpuwl/out, /root/gpuwl/bin, and /root/gpuwl/k8s, and create the namespace gpu-vm. Attach the label nvidia.com/gpu.workload.config to the three nodes — container on lab-node-0, vm-passthrough on lab-node-1, and vm-vgpu on lab-node-2. And write three lines in /root/gpuwl/out/01-nodes.txt — each line is <노드이름> <라벨값> (the placeholders are the node name and the label value), in ascending order of node name.
  2. Write five lines in /root/gpuwl/out/operands.txt. The first three lines are <라벨값>=<오퍼랜드 목록> (the placeholders are the label value and the operand list), and the list is joined by commas (order does not matter) — container is datacenter-driver, container-toolkit, device-plugin, and dcgm-exporter; vm-passthrough is vfio-manager and sandbox-device-plugin; and vm-vgpu is vgpu-manager, vgpu-device-manager, and sandbox-device-plugin. On the fourth line, write in NO_LABEL= the value the Operator assumes when there is no label, and on the fifth line, write in ENABLE_FLAG= the name of the ClusterPolicy flag that makes this label get used.
  3. In both status.capacity and status.allocatable of the three nodes, put each one's resource — nvidia.com/gpu as "4" on lab-node-0, nvidia.com/GA102GL_A10 as "2" on lab-node-1, and nvidia.com/NVIDIA_A10-12Q as "4" on lab-node-2. And write three lines in /root/gpuwl/out/03-resources.txt — each line is <노드이름> <자원이름> <광고량> (the placeholders are the node name, the resource name, and the advertised amount), in ascending order of node name.
  4. In /root/gpuwl/k8s/pod-container.yaml, write the Pod job-container — namespace gpu-vm, container name cuda, image nvcr.io/nvidia/cuda:12.4.1-base-ubuntu22.04, and nvidia.com/gpu: 1 in limits. Do not use a nodeSelector. After applying it, write in one line in /root/gpuwl/out/04-container.txt which node it came up on — NODE=<노드이름> (the placeholder is the node name).
  5. In /root/gpuwl/k8s/pod-passthrough.yaml, write the Pod vmi-passthrough — namespace gpu-vm, nvidia.com/GA102GL_A10: 1 in limits, and no nodeSelector. The rest is the same as step 4. After applying it, write two lines in /root/gpuwl/out/05-passthrough.txt — NODE= and RESOURCE= (the requested resource name as it is).
  6. In /root/gpuwl/k8s/pod-vgpu.yaml, write the Pod vmi-vgpu — namespace gpu-vm, nvidia.com/NVIDIA_A10-12Q: 1 in limits, and no nodeSelector. After applying it, write three lines in /root/gpuwl/out/06-placement.txt — where each of the three Pods made so far came up, as <파드이름> <노드이름> (the placeholders are the Pod name and the node name), in ascending order of Pod name.
  7. In /root/gpuwl/k8s/pod-cross.yaml, write the Pod job-cross — namespace gpu-vm, hang nvidia.com/gpu.workload.config: vm-passthrough as the nodeSelector, and require nvidia.com/gpu: 1 in limits. It amounts to asking a passthrough node for a container resource. After applying it, write three lines in /root/gpuwl/out/07-cross.txt — PHASE=, NODE= (none if empty), and MESSAGE= (the message of the PodScheduled condition).
  8. Save the result of kubectl get nodes -o json to /root/gpuwl/out/nodes.json, and create /root/gpuwl/bin/check-config.sh <노드JSON파일> (the placeholder is the node JSON file). It looks only at those nodes in the file that have the nvidia.com/gpu.workload.config label and judges whether that value and the nvidia.com/ resource name advertised agree — container must be nvidia.com/gpu, vm-vgpu must be a vGPU profile name (one ending in a number and a capital letter), and vm-passthrough must be a device model name that is neither of those. For each node that disagrees it prints one line MISMATCH=<노드이름> <이유> (the placeholders are the node name and the reason) and ends with 1, and if all agree it prints OK=<검사한 노드 수> (the placeholder is the number of nodes checked) and ends with 0. After creating it, run it on the dump you saved and save the output to /root/gpuwl/out/consistency.txt.

Notes

Write the use on three nodes

Create /root/gpuwl/out, /root/gpuwl/bin, and /root/gpuwl/k8s, and create the namespace gpu-vm. Attach the label nvidia.com/gpu.workload.config to the three nodes — container on lab-node-0, vm-passthrough on lab-node-1, and vm-vgpu on lab-node-2. And write three lines in /root/gpuwl/out/01-nodes.txt — each line is <노드이름> <라벨값> (the placeholders are the node name and the label value), in ascending order of node name.

This one label swaps wholesale the Operator's software that will come up on that node. Attach the label with kubectl label node <이름> --overwrite <키>=<값> (the placeholders are the name, the key, and the value). There are only three values and even if you make a typo there is no warning — so the habit of checking after attaching is needed. You can see it on one screen with kubectl get nodes -L <키> (the placeholder is the key).

Lay out in a table what comes up for each label value

Write five lines in /root/gpuwl/out/operands.txt. The first three lines are <라벨값>=<오퍼랜드 목록> (the placeholders are the label value and the operand list), and the list is joined by commas (order does not matter) — container is datacenter-driver, container-toolkit, device-plugin, and dcgm-exporter; vm-passthrough is vfio-manager and sandbox-device-plugin; and vm-vgpu is vgpu-manager, vgpu-device-manager, and sandbox-device-plugin. On the fourth line, write in NO_LABEL= the value the Operator assumes when there is no label, and on the fifth line, write in ENABLE_FLAG= the name of the ClusterPolicy flag that makes this label get used.

If you put the three lines side by side, you see that the only common denominator is sandbox-device-plugin — the two VM-side uses advertise devices the same way and differ in how they prepare the driver. The fifth line is the most important. If that flag is off (the default is off), however exactly you attach the label it is not read, and every node is prepared for containers. The flag name is two words joined by a dot.

Advertise a different resource name for each use

In both status.capacity and status.allocatable of the three nodes, put each one's resource — nvidia.com/gpu as "4" on lab-node-0, nvidia.com/GA102GL_A10 as "2" on lab-node-1, and nvidia.com/NVIDIA_A10-12Q as "4" on lab-node-2. And write three lines in /root/gpuwl/out/03-resources.txt — each line is <노드이름> <자원이름> <광고량> (the placeholders are the node name, the resource name, and the advertised amount), in ascending order of node name.

In a real cluster, a different device plugin per node fills this spot — the Kubernetes device plugin on the container node and the sandbox device plugin on the two VM-side nodes. The passthrough resource name comes from the PCI device model, and the vGPU resource name comes from the profile — the rule that only the latter ends in a number and a capital letter becomes the basis of judgment in the last step. You need not escape the dot and slash of the resource name in a JSON patch. You escape the dot only when reading with jsonpath.

A container workload goes to the container node

In /root/gpuwl/k8s/pod-container.yaml, write the Pod job-container — namespace gpu-vm, container name cuda, image nvcr.io/nvidia/cuda:12.4.1-base-ubuntu22.04, and nvidia.com/gpu: 1 in limits. Do not use a nodeSelector. After applying it, write in one line in /root/gpuwl/out/04-container.txt which node it came up on — NODE=<노드이름> (the placeholder is the node name).

Though you gave no condition at all, there is only one node it can go to. The resource name itself chose the node — the point of this whole lab is contained in this one line. Write an extended resource only in limits. Because of the rule that requests and limits must be equal, limits alone is enough.

A passthrough request calls the device model name

In /root/gpuwl/k8s/pod-passthrough.yaml, write the Pod vmi-passthrough — namespace gpu-vm, nvidia.com/GA102GL_A10: 1 in limits, and no nodeSelector. The rest is the same as step 4. After applying it, write two lines in /root/gpuwl/out/05-passthrough.txt — NODE= and RESOURCE= (the requested resource name as it is).

In a real cluster a VirtualMachineInstance comes at this spot, and you write the resource in spec.domain.devices.gpus[].deviceName. This Pod has no KubeVirt so it cannot bring up a VM, but what happens at the scheduling stage is the same — because a VM too is in the end run on its behalf by a Pod that requires that resource. The resource name comes from the PCI device model. Get even one character wrong and that node disappears from the candidates.

A vGPU request calls the profile name

In /root/gpuwl/k8s/pod-vgpu.yaml, write the Pod vmi-vgpu — namespace gpu-vm, nvidia.com/NVIDIA_A10-12Q: 1 in limits, and no nodeSelector. After applying it, write three lines in /root/gpuwl/out/06-placement.txt — where each of the three Pods made so far came up, as <파드이름> <노드이름> (the placeholders are the Pod name and the node name), in ascending order of Pod name.

If you put the three lines side by side, you see at a glance that a single resource name decided the whole placement. Though you gave no nodeSelector to any Pod, the three each went to a different node. A vGPU profile name ends in a number and a capital letter (for example, 12Q). That form is the point at which it is distinguished from a passthrough name, and it is also the signal the step 8 checker will use. You can extract the placement all at once with kubectl get pods -n <ns> -o custom-columns= (the placeholder is the namespace).

If you call another's resource, it cannot go anywhere

In /root/gpuwl/k8s/pod-cross.yaml, write the Pod job-cross — namespace gpu-vm, hang nvidia.com/gpu.workload.config: vm-passthrough as the nodeSelector, and require nvidia.com/gpu: 1 in limits. It amounts to asking a passthrough node for a container resource. After applying it, write three lines in /root/gpuwl/out/07-cross.txt — PHASE=, NODE= (none if empty), and MESSAGE= (the message of the PodScheduled condition).

That node advertises two whole cards and the slots are empty. Yet it cannot go — because the scheduler only compares resource names as strings. If a cluster splits its inventory by use, then even when the queue on one side is long, the slots on the other side stay empty. This is the first constraint you must plan for when introducing sandbox workloads. Take note of what the message says — it says "resources are short".

A checker that judges whether labels and resource advertisements agree

Save the result of kubectl get nodes -o json to /root/gpuwl/out/nodes.json, and create /root/gpuwl/bin/check-config.sh <노드JSON파일> (the placeholder is the node JSON file). It looks only at those nodes in the file that have the nvidia.com/gpu.workload.config label and judges whether that value and the nvidia.com/ resource name advertised agree — container must be nvidia.com/gpu, vm-vgpu must be a vGPU profile name (one ending in a number and a capital letter), and vm-passthrough must be a device model name that is neither of those. For each node that disagrees it prints one line MISMATCH=<노드이름> <이유> (the placeholders are the node name and the reason) and ends with 1, and if all agree it prints OK=<검사한 노드 수> (the placeholder is the number of nodes checked) and ends with 0. After creating it, run it on the dump you saved and save the output to /root/gpuwl/out/consistency.txt.

This checker does not attach to the cluster and reads only one file. That way it can be used on a dump someone else sent you and in a pre-deployment review. The core of the judgment is the shape of the name — a vGPU profile ends after a hyphen with digits and one capital letter, like -12Q and -4Q. A node advertising several kinds of nvidia resources is also a disagreement. A worker node runs only one kind of GPU workload. A resource advertised as "0" must be treated as absent. The grader also runs it on a dump deliberately made to disagree.