TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

workload.config — one label decides what a node is

Continue in TT Lab

In one line

The GPU Operator looks at the value of the nvidia.com/gpu.workload.config label (container, vm-passthrough, vm-vgpu) and swaps out wholesale which operands to bring up on that node, and as a result even the resource name the node advertises changes. A container node advertises nvidia.com/gpu, a passthrough node advertises a device model name such as nvidia.com/GA102GL_A10, and a vGPU node advertises a profile name such as nvidia.com/NVIDIA_A10-12Q.

Why this was needed

Not all workloads that use GPUs are containers. There are in fact many cases where a license requires a particular OS, a Windows driver is needed, or a legacy simulation lives entirely inside a virtual machine. Such organizations long operated their clusters split in two — one Kubernetes for containers and one hypervisor for virtual machines. If you split the same GPU inventory between two places, you cannot lend to one side even when the other is idle.

KubeVirt removed that wall by making virtual machines Kubernetes objects, and the GPU Operator became able, on top of that, to prepare worker nodes separately for two kinds of use. The official documentation summarizes the difference in one line — containers need the data center driver, GPU passthrough needs the vfio-pci driver, and vGPU needs the NVIDIA vGPU Manager. That the required drivers differ is the root of every difference.

One label swaps a node's software wholesale

Attaching a label to the node is all there is to it.

kubectl label node <노드이름> --overwrite nvidia.com/gpu.workload.config=vm-vgpu

Then the operands the Operator brings up on that node change.

Label value What comes up on that node What it does
container Data center driver · container toolkit · Kubernetes device plugin · DCGM and DCGM Exporter Makes containers able to use the GPU and emits metrics
vm-passthrough VFIO Manager · sandbox device plugin Loads vfio-pci and binds it to all GPUs on that node, and advertises the passthrough GPUs to kubelet
vm-vgpu vGPU Manager · vGPU Device Manager · sandbox device plugin Installs the vGPU driver, creates vGPU devices, and advertises them to kubelet

There are two sentences you must hold on to here.

First, without the label the default is container. The Operator prepares a node without the label for containers. To change the default, you edit sandboxWorkloads.defaultWorkload of the ClusterPolicy.

Second, this label is used only when sandboxWorkloads.enabled is on. That flag is off by default, and when it is off, every node is prepared for containers and the label is not read at all. This is the spot where people attach only the label and stare for hours at why nothing changes. The label is neither a syntax error nor does it raise a warning.

And there is one more constraint the official documentation nails down. A worker node runs only one kind of GPU workload. You cannot mix containers and passthrough VMs on one node. That is because vfio-pci takes all of that node's GPUs, so the host's data center driver must let go of those cards. It means inventory planning hardens at the node level, and this is the constraint felt most strongly in practice.

The resource name changes

What a container node advertises is the nvidia.com/gpu we know. A sandbox node is different. The output example in the official documentation says it as it is.

# 패스스루 노드
$ kubectl get node <노드> -o json | jq '.status.allocatable | with_entries(select(.key | startswith("nvidia.com/")))'
{ "nvidia.com/GA102GL_A10": "1" }

# vGPU 노드 (기본 설정: 카드마다 절반 크기 Q 프로파일 두 개)
{ "nvidia.com/NVIDIA_A10-12Q": "4" }

The former comes from the PCI device model name, and the latter from the vGPU profile name. If you apply the default configuration to two A10 cards, two 12Q instances are created per card, for a total of 4. Here you can change the profile with the node label nvidia.com/vgpu.config — if you choose A10-4Q, six per card are advertised, which is 12 for two cards. Though the hardware is the same, a single label changes the advertised amount and the resource name together.

The consequence of the names being different is simple and harsh. A Pod that requires nvidia.com/gpu can never go to a passthrough node. Even if slots remain, to the scheduler that node is a node that does not have that resource. The reverse is the same. If a cluster splits its card inventory, then even when the queue on one side is long, the slots on the other side stay empty.

What it looks like to plug a card into a VM

Just because a device was advertised does not mean a VM can use it right away. You must separately write an allowlist on the KubeVirt side — this is the step people miss most. Under permittedHostDevices of the KubeVirt custom resource, you write pciHostDevices for passthrough and mediatedDevices for vGPU, and attach externalResourceProvider: true to both. That flag means "this resource was not created by us but is advertised by an external device plugin (here, the sandbox device plugin)".

spec:
  configuration:
    developerConfiguration:
      featureGates:
        - GPU
        - DisableMDEVConfiguration
    permittedHostDevices:
      pciHostDevices:
        - externalResourceProvider: true
          pciVendorSelector: 10DE:2236
          resourceName: nvidia.com/GA102GL_A10

Only after that can you request the card from the VM side. The place is different from a Pod's resources.limits — a VirtualMachineInstance writes it in spec.domain.devices.gpus.

spec:
  domain:
    devices:
      gpus:
        - deviceName: nvidia.com/GA102GL_A10
          name: gpu1

deviceName is the resource name and name is the alias by which that device is called inside the VM. The outward shape is different, but what happens at the scheduling stage is the same as for a Pod — the Pod that runs the VM on its behalf requires that resource and goes to a node that advertises that resource.

Finally, one fact the official documentation states clearly. The GPU Operator does not automate the installation of the NVIDIA driver inside the VM. Its job goes as far as plugging the card into the VM, and the driver inside the guest OS is the responsibility of whoever builds the VM image.

What it looks like in the field

First, BIOS and kernel parameters are prerequisites. Virtualization extensions and IOMMU must be turned on in the BIOS, and the kernel command line must have intel_iommu=on or amd_iommu=on. To use vGPU with Ampere and later cards, you must also turn on SR-IOV in the BIOS. These are changes that need a reboot, so if you do not do them at the image stage before putting a node into the cluster, you must later fix them by emptying nodes one at a time.

Second, moving a node between uses is not zero-downtime. Changing the label splits the operands and changes the driver. The official documentation says that when you change the vGPU configuration, you must first shut down or move the VMs that were running on that node. The plan to run inventory flexibly meets reality at this point — moving by the day works, but by the minute does not.

Third, because the resource name is tied to the card model, manifests get tied to hardware. A VM manifest that requires nvidia.com/GA102GL_A10 stays Pending forever in a cluster with no A10. This contrasts with how nvidia.com/gpu on the container side was unrelated to model. So organizations that use sandbox workloads do not write the model name directly in manifests but gather it in one place with Helm values or a kustomize patch.

Fourth, the honest limit of this environment. The lab cluster has neither KubeVirt nor vfio-pci and cannot bring up VMs. So in the next lab you set up the VM request as a Pod that requires the same resource name. Because what happens at the scheduling stage is in fact the same (both are Pods requiring an extended resource), where it goes and where it gets blocked is learned with the real thing. However, the latter part where the VM actually boots and holds the card cannot be confirmed, and that part is replaced by the shape of the VMI manifest above.

Reference documents

What you will do in the next lab

You label three nodes container, vm-passthrough, and vm-vgpu, and set up in the node status the resource name that fits each label. You organize into a table file which operands must come up for each label value, bring up three workloads (a container Pod, a passthrough request, and a vGPU request), and confirm with the real scheduler that each goes to its own node and that one requiring another's resource cannot come up anywhere. Finally, you build a checker that takes a node dump and judges "do this label and this resource advertisement agree with each other", and see whether it catches a node deliberately made to disagree.