The Resource Gate — Pending, Decided by Advertising, Request Rules and Taints
Goal
You walk through, to the end, the resource side of the two gates a GPU Pod must pass. You make the node advertise the resource, see the two rules for requesting extended resources actually block things, and go as far as reading whether the cause of the same Pending is a taint or resources.
Why it matters
If, in the report "the GPU Pod does not come up", you cannot decide the strand in the first 3 minutes, you spend hours in the wrong place. If the Pod has not even been assigned to a node, that is the scheduler stage, and what the scheduler sees is not the devices but the number written in the node status. What writes that number is the device plugin, and without that number, the Pod waits forever even if the device is perfectly plugged in.
Extended resources have two more rules. requests and limits must be equal, and the value must be an integer. That is because overcommit is impossible and a device cannot be split. Both rules are checked by the apiserver, so if you break them, the Pod is not created at all. Finally, GPU nodes usually have a taint, so even when resources are plentiful, a Pod is not scheduled without a toleration. These three are the whole of the resource gate.
Environment
This Pod has neither a GPU nor a device plugin. Instead, there is a real control plane brought up by kwok, so if you write a number in the node status, the real scheduler sees it and judges. The validation of extended resources is also done by the real apiserver. The working directory is /root/gpuwhy, with manifests in k8s/ and outputs in out/.
Steps
- Create the
gpu-whynamespace and attach four discovery labels tolab-node-0. - Advertise
nvidia.com/gpu: 2in that node's status and save it toout/advertised.txt. - Create two manifests that break the rules and save the rejection messages to
out/rejected.txt. - Create the
trainerPod without conditions and see the node being decided by resources alone. - Put a taint on the node and save the reason
cpu-onlywaits toout/tainted.txt. - Pass both gates together with
trainer2, which has a toleration attached. - Hand over the slots with
trainer3and save toout/exhausted.txtthe change of the waiting reason to resources. - Summarize in seven lines in
out/gate-report.txt.
Notes
- status is a subresource. Use
kubectl patch node <이름> --subresource=status(the placeholder is the name), and in a JSON pointer, escape a slash as~1. - If you write only
capacityand leave outallocatable, there are no slots for the scheduler to use. It is a common mistake, so if the Pod keeps waiting even after step 3, look here first. - Steps 5 and 7 are both Pending. Put the two files side by side and compare how the messages differ. That difference decides where to look next.
- Do not get past the taint by removing it. The reason you reserved the GPU node disappears.
Attach discovery labels to the GPU node
Create the gpu-why namespace and attach four labels to lab-node-0. They are nvidia.com/gpu.present=true, nvidia.com/gpu.product=NVIDIA-A100-SXM4-40GB, nvidia.com/gpu.count=2, and nvidia.com/gpu.deploy.device-plugin=true. Do not attach them to the other two nodes.
In a real cluster, the gpu-feature-discovery DaemonSet reads the devices and attaches these labels. The Operator gates which DaemonSet to bring up on which node with the nvidia.com/gpu.deploy.* labels. Here you make only that result by hand. Use kubectl label node <이름> <키>=<값> --overwrite (the placeholders are the name, the key, and the value), and check with kubectl get node --show-labels.
Advertise the extended resource in the node status
Write nvidia.com/gpu as 2 in the status of lab-node-0. You must put it in both capacity and allocatable. Then save the advertised amounts of the three nodes to /root/gpuwhy/out/advertised.txt.
An extended resource is not a value kubelet counts by itself but a number in which what the device plugin reported is written in the node status. Here you write that spot directly. status is a separate subresource, so you edit it with kubectl patch node <이름> --subresource=status --type=json -p '[...]' (the placeholder is the name). In a JSON pointer, a slash must be escaped as ~1, so the path becomes /status/capacity/nvidia.com~1gpu. If you write only capacity, no slots appear for the scheduler to use.
Confirm what the two rules of extended resources block
Create two Pod manifests that will be rejected. /root/gpuwhy/k8s/bad-mismatch.yaml writes different requests and limits for nvidia.com/gpu, and /root/gpuwhy/k8s/bad-fraction.yaml requests a fraction. Apply both and save the rejection messages to /root/gpuwhy/out/rejected.txt.
Extended resources do not allow overcommit, so requests and limits must be equal, and the value must be an integer. That is because devices are allocated in units that cannot be split. Both rules are checked by the apiserver, so the Pod is not created at all. The command failing is the right answer, and the message goes to standard error, so you must capture it with 2>&1 for it to remain in the file.
Even without conditions, resources decide the node
Create the Pod trainer in the gpu-why namespace with /root/gpuwhy/k8s/trainer.yaml. Write nvidia.com/gpu as 1 in limits only, and do not attach a nodeSelector.
For an extended resource, if you write only limits, requests is filled in automatically with the same value. After creating it, check directly with kubectl -n gpu-why get pod trainer -o yaml that requests was filled in. The point of this step is that it goes to a node that advertises a GPU even though you did not specify a node. What the scheduler sees is not devices but numbers.
Even with resources to spare, a taint blocks
Put the taint nvidia.com/gpu=present:NoSchedule on lab-node-0, and create, with /root/gpuwhy/k8s/cpu-only.yaml, a Pod cpu-only that does not use the GPU. That Pod must aim at the GPU node with the nodeSelector nvidia.com/gpu.present: "true". When it goes into the waiting state, save the scheduling condition message to /root/gpuwhy/out/tainted.txt.
GPU nodes are expensive per unit, so if ordinary workloads fill the CPU slots, there is no room left for the GPU Pods that actually need it. A taint is the mechanism that reserves that node for GPU use only. Apply it with kubectl taint node <이름> <키>=<값>:<효과> (the placeholders are the name, the key, the value, and the effect), and extract the waiting reason with kubectl -n gpu-why get pod cpu-only -o jsonpath='{.status.conditions[?(@.type=="PodScheduled")].message}'. Also note that Pods that were already running are not evicted by NoSchedule.
Attach a toleration and pass both gates together
Create the Pod trainer2 with /root/gpuwhy/k8s/trainer2.yaml. Put in together a toleration that tolerates the taint you applied earlier and a request for 1 card of nvidia.com/gpu. Do not remove the taint.
The key, value, and effect of the toleration must match the node's taint. The check point is that on the same node cpu-only keeps waiting and only trainer2 gets in. If you pass by removing the taint, the reason you reserved the GPU node disappears.
When the slots are full, even the same Pod waits
Create one more Pod trainer3, identical to trainer2, with /root/gpuwhy/k8s/trainer3.yaml. When it goes into the waiting state, save the scheduling condition message to /root/gpuwhy/out/exhausted.txt.
The node advertised 2 slots, and the first two Pods have already used them all. You must leave the toleration as it is for it to become clear that this time the waiting reason is resources, not the taint. Confirm that the scheduler's message changes to Insufficient nvidia.com/gpu. Even the same Pending has a different place to look if the reason is different.
Summarize the gate in numbers and a single word
Write seven lines in /root/gpuwhy/out/gate-report.txt. They are ALLOCATABLE, USED, PENDING, TAINT, REQUESTS_EQUAL_LIMITS, FRACTIONAL, and SCHEDULER_SEES.
Do not guess the three numbers; count them in the cluster and write them. USED is the sum of the nvidia.com/gpu requests of the Pods assigned to lab-node-0, and PENDING is the number of Pods waiting in the gpu-why namespace. For TAINT, write the one you applied in step 5 in the form 키=값:효과 (the fields are the key, the value, and the effect), and for the other three, write in a single word the rules you confirmed in steps 3 and 4. SCHEDULER_SEES is the answer to whether the scheduler sees devices or numbers.