ClusterPolicy, RuntimeClass and Resource Advertising
In one line
For a GPU Pod to come up, it must pass two gates that have no relationship to each other. The nvidia.com/gpu resource the scheduler looks at, and the RuntimeClass handler the node's runtime looks at. If you have only one, the Pod does not even tell you which of the two blocked it.
Why this distinction is needed
When you receive the report "the GPU Pod does not come up", the symptom is exactly one of two things.
- The Pod does not move from
Pending. → It is the resource side. No node has advertisednvidia.com/gpu, or the slots are full. - The Pod attached to a node but dies at
ContainerCreating. → It is the runtime side. The handler is not registered in that node's containerd.
If you cannot tell these two apart, you dig through the device plugin logs and spend hours on what is really a containerd problem. The reverse is the same. If you know why the two gates are separate, you can decide the strand in the first 3 minutes.
How it works
The resource side. A GPU is not a resource that kubelet counts by itself, like cpu and memory. When the device plugin opens the devices on the node and hands the list to kubelet, kubelet writes that count as an extended resource in status.capacity and status.allocatable of the node object. What the scheduler sees is not the devices but that number. So scheduling holds even with only a number and no device (that is the spot this course's lab stands on), and conversely, if the devices are fine but the plugin is dead, the scheduler judges that the node has no GPU.
Extended resources have one more rule. You cannot give requests and limits separately. The value written in limits becomes the requests, and it must be an integer. That is because you cannot give a GPU 0.5 of a card.
The runtime side. A RuntimeClass is a cluster-scoped object that chooses "with which runtime handler to create this Pod".
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: nvidia
handler: nvidia # 이 문자열이 containerd 설정의 runtimes.<이름> 과 맞아야 한다
When kubelet creates a Pod, it carries this handler string in the CRI request, and containerd looks for the same name in the runtimes table of its own configuration. If it is not there, container creation is refused. What matters here is that the API server does not validate this name. A RuntimeClass can be created freely with kubectl apply, and even if it points to a handler that does not exist, nobody warns you. The mismatch shows up at the moment the Pod is created, and only on that node.
The relationship of the two gates in one line is this.
스케줄러 : nvidia.com/gpu 숫자를 보고 "어느 노드로 보낼까" 를 정한다
kubelet : RuntimeClass handler 를 containerd 에 넘겨 "어떻게 만들까" 를 정한다
What it looks like in the field
First, there are many clusters where you need not write runtimeClassName. That is because it is common for the toolkit to change containerd's default_runtime_name to nvidia outright. It is convenient but has a price — even Pods that do not use the GPU go through the nvidia runtime, and if that configuration breaks, it is not the GPU Pods but all Pods in the cluster that cannot come up. The scale of the accident changes.
Second, it is moving to CDI. The Container Device Interface standardizes "how to put a device into a container" as a JSON/YAML spec independent of the runtime type. With CDI, you can inject by a device name such as nvidia.com/gpu=all without an nvidia-specific runtime binary. However, the way to turn on CDI support on the containerd side differs by version, so the field is now a mix of the two methods.
Third, the cause of Pending is not only resources. GPU nodes usually have a taint such as nvidia.com/gpu=present:NoSchedule, which keeps ordinary workloads from occupying the expensive nodes. If a GPU Pod has no toleration, it is not scheduled even when resources are plentiful. This is why the last event in kubectl describe pod is always the first sentence to read.
What you will do in the next lab
In the next lab, you walk through the resource side of the two gates, in eight steps. You write an extended resource in the node status and have the scheduler judge it, see the requests and limits rules get blocked at the apiserver, and read from the message whether the cause of the same Pending is a taint or resources. If you can tell the strands apart exactly here, the incident stories that follow are read much faster.