Same message, different cause — a triage table for GPU failures
In one line
When you receive the report "the GPU Pod doesn't come up", first split it into two strands by whether the Pod object exists or not, then if it does, split again by whether it is Pending or Running, and then cross-check the node status against the Pod spec to narrow it to one of seven causes. The sentence the scheduler leaves comes out identical down to the letter from three different causes, so reading the message is not enough.
Why this was needed
The inquiry that comes most often in GPU cluster operations is always the same one sentence. "My job isn't running." But the actual states behind this one sentence differ at least this much.
- You tried to create a Pod and
kubectl applyraised an error. The Pod does not exist. - The Pod was created but has been
Pendingfor hours. - The Pod is
Runningbutnvidia-smiinside it cannot find the device.
The three have entirely different places to investigate. The first is admission (quota, RuntimeClass, policy), the second is the scheduler, and the third is the Pod spec and the node's runtime. But the reporter does not distinguish these three for you. So the receiving side must have a procedure that decides the strand in the first 3 minutes. Without that procedure, you dig through the device plugin logs for an hour and only then learn "oh, the Pod was never created in the first place".
There is a circumstance that is one step worse. The Kubernetes scheduler tends to leave failure reasons kindly, but for extended resources that kindness works in the direction of hiding the cause. The sentences that come out in the three situations below are the same.
- That node is not advertising
nvidia.com/gpuat all (the device plugin died or does not exist). capacityhas 4 cards butallocatableis 0 (driver validation failed, the node is failing to hand out GPUs).- The advertisement is fine and there are slots, but other Pods already hold them all.
In all three cases it is Insufficient nvidia.com/gpu. The causes are entirely different — "revive the plugin", "fix the node", "wait for a slot or add more" — but the message cannot tell them apart. This one fact is the reason this module exists.
The first two questions that split the symptoms
Question 1 — does the Pod object exist?
kubectl get pod <이름> -n <네임스페이스> -o json
If it does not, the scheduler has nothing to do with this case. An object comes to exist only by passing the API server's admission, and the two typical things admission blocks regarding GPUs are exceeding a ResourceQuota and referencing a RuntimeClass that does not exist. In both, the error is printed only on the terminal of the person who ran kubectl apply and no trace remains in the cluster. If the reporter has closed that screen, the fastest way is to apply the same manifest again and reproduce the error.
Question 2 — is spec.nodeName empty?
kubectl get pod <이름> -n <네임스페이스> -o jsonpath='{.spec.nodeName}'
If it is empty, it is a scheduling problem. If it is filled and there is still a problem, it is a problem on the node (the runtime handler, the driver, a manifest that did not make a request). The reason you look at this field rather than the STATUS column of kubectl get pod is that the single word Pending covers both "could not choose a node" and "chose a node but cannot create the container".
In which column the cause shows up
| Cause | Where the decisive evidence is | The command to look |
|---|---|---|
| Quota exceeded | The Pod does not exist. The error wording at creation | kubectl get resourcequota -n <ns> -o json |
| No RuntimeClass | The Pod does not exist. The error wording at creation | kubectl get runtimeclass |
| No request made | nvidia.com/gpu is absent from spec.containers[].resources.limits |
kubectl get pod -o json |
| Label mismatch | There are 0 nodes satisfying spec.nodeSelector |
kubectl get nodes --show-labels |
| Taint | The Pod's tolerations cannot tolerate the candidate node's spec.taints |
kubectl get node -o json |
| Resource not advertised | The resource name itself is absent from the candidate node's status.capacity |
kubectl get node -o json |
| allocatable 0 | capacity is positive but status.allocatable is 0 |
Another column of the same output |
| Slots exhausted | allocatable is positive but the sum of Pod requests on that node equals it | Sum up kubectl get pods -A -o json |
The usefulness of this table lies not in "what to look at" but in "in what order to look." The order is from top to bottom. If an earlier one is true, you need not look at the later ones, and if you do not keep the order, you reach wrong conclusions. For example, a Pod whose nodeSelector matches no node has no taint or resource to check in the first place. Yet if you first take an action such as "I removed the taint and it still doesn't come up", you only break a perfectly good cluster configuration.
The last row (slots exhausted) is the only one that requires computation. You cannot tell from the node object alone; you must gather the Pods placed on that node and add up their GPU requests. There is one trap here. Pods that ended as Succeeded or Failed do not hold slots, so they must be excluded from the sum. If you do not exclude them, you wrongly judge a node that is actually empty as "full".
Circumstances that arise because it is an extended resource
nvidia.com/gpu is, unlike cpu and memory, an extended resource that kubelet cannot count by itself. The official documentation nails down the rules of extended resources in two. Overcommit is not supported, requests and limits must be equal, and the value must be an integer. So when reading a GPU request, looking only at limits is enough. Conversely, values such as half a card or 1.5 cards cannot exist, and there is no behavior of receiving only as much as is available. A request is either satisfied in full or the Pod waits, one or the other.
And an extended resource holds regardless of "whether the device exists". If only a number is written in the node status, the scheduler sends Pods to that node. Conversely, even if eight cards are plugged in, if there is no number in the status, that node is a node without a GPU to the scheduler. When investigating, you must hold on to the fact that not the physical fact but the node object is the scheduler's only truth.
What it looks like in the field
First, the habit of asserting the cause after reading only kubectl describe is the most expensive. As seen earlier, three causes give the same sentence. In particular, there was an organization that hung a script that automatically adds nodes when it sees "Insufficient nvidia.com/gpu" at night, and the real cause was that one node's allocatable had dropped to 0. Even when they added nodes, that node stayed at 0, and only the cost increased.
Second, when people judge, they cannot keep the order. So it is better to harden this judgment into a tool. The input is one Pod and the output is one word for the cause. With such a tool, the first responder can answer "this is the quota" on the spot where the report was received, and the criteria of judgment do not vary from person to person. What this module's lab makes is exactly that tool.
Third, the output must be one word to be useful. If you emit a sentence, a person has to read it again, and it cannot be hooked into automation. If you choose a vocabulary in which the next action is fixed to one, such as no-gpu-node, allocatable-zero, and exhausted, it directly becomes an alert routing key. Deciding the cause vocabulary is in fact deciding the operating procedure.
Fourth, the honest limit of this environment. The lab cluster has neither a GPU nor a device plugin. So the state "the device plugin died" is made not by killing the plugin but by not writing the resource in the node status. This is not an imitation but the same state — because even in a real failure, what the scheduler looks at is only the status. However, the latter half of "the Pod is Running but the device is not visible inside the container" cannot be confirmed because containers do not actually run. That strand is judged only as far as the fact that the Pod spec has no resource request.
Reference documents
- Advertising extended resources for a node: https://kubernetes.io/docs/tasks/administer-cluster/extended-resource-node/
- Managing resources for containers (the extended resource rules): https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/
- Taints and tolerations: https://kubernetes.io/docs/concepts/scheduling-eviction/taint-and-toleration/
- Resource quotas: https://kubernetes.io/docs/concepts/policy/resource-quotas/
- Debugging Pods: https://kubernetes.io/docs/tasks/debug/debug-application/debug-pods/
- GPU Operator troubleshooting: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/troubleshooting.html
What you will do in the next lab
On the real scheduler brought up by kwok, you set up the seven states one by one. You set up a node with a healthy advertisement, a node with no resource name at all, and a node with only capacity and allocatable at 0, attach a Pod to each, and see with your own eyes that the three causes give the same sentence. You add a Pod that differs only in taint and a Pod that differs only in nodeSelector, and experience the two cases rejected at admission so that not even the object comes to exist (RuntimeClass and quota). Finally, you build a classifier that answers the cause in one word looking only at kubectl get -o json, and run it on all seven Pods.