TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

GPU triage — telling different causes apart behind one message

Continue in TT Lab

Goal

You set up on a real scheduler the seven states in which a GPU Pod gets blocked, and build a classifier that answers the cause in one word by looking only at objects.

Why it matters

Behind the single report "the GPU Pod doesn't come up" there are half a dozen real causes, and the places to investigate are all different. On top of that, three of them — resource not advertised, allocatable 0, and slots exhausted — make the scheduler give exactly the same sentence. So reading the message is not enough; you must cross-check objects in the order nodeSelector → taint → capacity → allocatable → remaining slots for the cause to narrow to one. The reason the order matters is that if an earlier step is true, there is nothing to check at the later steps — if you break the order, you only break the cluster, for example by removing a perfectly good taint. If a person does this judgment every time, the criteria waver, so you harden it into a tool at the end.

Steps

  1. Create the working directories /root/gputri/cases, /root/gputri/out, and /root/gputri/bin, and create the namespace gpu-triage. In both status.capacity and status.allocatable of lab-node-0, put nvidia.com/gpu as "2", and apply the taint nvidia.com/gpu=present:NoSchedule. In /root/gputri/cases/case-ok.yaml, write the Pod case-ok — namespace gpu-triage, nodeSelector kubernetes.io/hostname: lab-node-0, a toleration that tolerates that taint, container name cuda, image nvcr.io/nvidia/cuda:12.4.1-base-ubuntu22.04, and nvidia.com/gpu: 1 in limits. Apply it, and when it is Running, write two lines in /root/gputri/out/01-ok.txt — NODE= and PHASE=.
  2. Leave lab-node-1 advertising nothing (it is a node whose device plugin died). In /root/gputri/cases/case-nogpunode.yaml, write the Pod case-nogpunode — pin it to lab-node-1 with nodeSelector, keep the toleration the same as in step 1, and require nvidia.com/gpu: 1. After applying it, confirm that it is Pending, and save the reason and message of the PodScheduled condition to /root/gputri/out/02-nogpunode.txt as two lines, REASON= and MESSAGE=.
  3. In lab-node-2, put nvidia.com/gpu as "4" in status.capacity and as "0" in status.allocatable (a state in which driver validation failed and the node cannot hand out GPUs). In /root/gputri/cases/case-alloc0.yaml, write the Pod case-alloc0 — pin it to lab-node-2 and the rest is the same as step 2. After applying it, write three lines in /root/gputri/out/03-alloc0.txt — CAPACITY=, ALLOCATABLE=, and MESSAGE=. Compare with the message from the previous step.
  4. In /root/gputri/cases/case-taint.yaml, write the Pod case-taint — pin it to lab-node-0 and require nvidia.com/gpu: 1, but do not put in a toleration. The rest is the same as the step 1 Pod. After applying it, write two lines in /root/gputri/out/04-taint.txt — TAINT=nvidia.com/gpu=present:NoSchedule and MESSAGE=. Check from the message why it is blocked even though resources remain.
  5. In /root/gputri/cases/case-label.yaml, write the Pod case-label — set the nodeSelector to nvidia.com/gpu.product: NVIDIA-A100-SXM4-40GB (no node has this label), and leave the toleration and the nvidia.com/gpu: 1 request as they are. After applying it, write two lines in /root/gputri/out/05-label.txt — for MATCHING_NODES=, the number of nodes that actually have that label, and for MESSAGE=, the condition message.
  6. In /root/gputri/cases/case-exhaust.yaml, write the Pod case-exhaust — pin it to lab-node-0, put in the toleration, and require nvidia.com/gpu: 2. That node advertises 2 cards, but case-ok from step 1 already holds one. After applying it, write three lines in /root/gputri/out/06-exhaust.txt — ALLOCATABLE= (the amount that node advertised), ALLOCATED= (the sum of the GPU requests of the Pods placed on that node), and REQUESTED=2.
  7. First, in /root/gputri/cases/case-norequest.yaml, write the Pod case-norequest — pin it to lab-node-0, put in the toleration, but do not put in a resource request at all. When applied, it comes up. Next, in /root/gputri/cases/case-rtc.yaml, write a Pod case-rtc carrying runtimeClassName: nvidia and apply it, and save the output, including standard error, to /root/gputri/out/07-runtimeclass.txt (do not create the RuntimeClass). Finally, create the namespace gpu-quota, write a ResourceQuota gpu-quota in /root/gputri/cases/quota.yaml that limits requests.nvidia.com/gpu to "1", then apply the Pod case-quota (requiring 3 GPUs) in /root/gputri/cases/case-quota.yaml and save the output to /root/gputri/out/07-quota.txt. Lastly, write three lines in /root/gputri/out/07-symptoms.txt — NOREQUEST=, RUNTIMECLASS=, and QUOTA=. On the latter two lines, write whether a Pod object was created as created or absent.
  8. Create /root/gputri/bin/triage.sh. When called as bash triage.sh <네임스페이스> <파드> (the placeholders are the namespace and the Pod), it prints only one word to standard output and ends. There are seven answers — no-request, ok, label-mismatch, taint, no-gpu-node, allocatable-zero, and exhausted. If the Pod does not exist, it prints a message to standard error and ends with 1. The judgment is made only from the objects given by kubectl get -o json, and the order is no request → already placed → nodeSelector → taint → capacity → allocatable → remaining slots. After creating it, run it on all seven Pods (case-ok, case-nogpunode, case-alloc0, case-taint, case-label, case-exhaust, and case-norequest) and write seven lines in /root/gputri/out/triage.txt in the form <파드이름> <원인> (the placeholders are the Pod name and the cause).

Notes

Set up one healthy GPU node

Create the working directories /root/gputri/cases, /root/gputri/out, and /root/gputri/bin, and create the namespace gpu-triage. In both status.capacity and status.allocatable of lab-node-0, put nvidia.com/gpu as "2", and apply the taint nvidia.com/gpu=present:NoSchedule. In /root/gputri/cases/case-ok.yaml, write the Pod case-ok — namespace gpu-triage, nodeSelector kubernetes.io/hostname: lab-node-0, a toleration that tolerates that taint, container name cuda, image nvcr.io/nvidia/cuda:12.4.1-base-ubuntu22.04, and nvidia.com/gpu: 1 in limits. Apply it, and when it is Running, write two lines in /root/gputri/out/01-ok.txt — NODE= and PHASE=.

kubelet cannot count extended resources by itself. In a real cluster the device plugin writes them in the node status, but here you write the same value in the same place directly with kubectl patch node <이름> --subresource=status --type=merge (the placeholder is the name). If you write only capacity, no slots appear for the scheduler to use — fill in both fields. Production GPU nodes almost always have a taint, so ordinary workloads do not drift in. That is why GPU Pods come with a toleration.

A node with no resource name at all

Leave lab-node-1 advertising nothing (it is a node whose device plugin died). In /root/gputri/cases/case-nogpunode.yaml, write the Pod case-nogpunode — pin it to lab-node-1 with nodeSelector, keep the toleration the same as in step 1, and require nvidia.com/gpu: 1. After applying it, confirm that it is Pending, and save the reason and message of the PodScheduled condition to /root/gputri/out/02-nogpunode.txt as two lines, REASON= and MESSAGE=.

You can extract the condition with kubectl get pod <이름> -n <ns> -o jsonpath='{.status.conditions[0].reason}' (the placeholder is the name). Take note of what is written in the message — in the next two steps, Pods with entirely different causes give the same sentence. Not having written nvidia.com/gpu on the node does not mean "there is no card" but "it looks absent to the scheduler".

A node that has capacity but cannot hand it out

In lab-node-2, put nvidia.com/gpu as "4" in status.capacity and as "0" in status.allocatable (a state in which driver validation failed and the node cannot hand out GPUs). In /root/gputri/cases/case-alloc0.yaml, write the Pod case-alloc0 — pin it to lab-node-2 and the rest is the same as step 2. After applying it, write three lines in /root/gputri/out/03-alloc0.txt — CAPACITY=, ALLOCATABLE=, and MESSAGE=. Compare with the message from the previous step.

capacity is "the amount this node has" and allocatable is "the amount it can hand out to the scheduler". It is normal for the two to diverge (the system reserve is that difference), and with GPUs it really happens that this difference opens up to 0. The two values are in different fields of the same kubectl get node -o json output. If you put the message side by side with the step 2 file, you see at a glance why this lab is needed.

A Pod missing only the toleration

In /root/gputri/cases/case-taint.yaml, write the Pod case-taint — pin it to lab-node-0 and require nvidia.com/gpu: 1, but do not put in a toleration. The rest is the same as the step 1 Pod. After applying it, write two lines in /root/gputri/out/04-taint.txt — TAINT=nvidia.com/gpu=present:NoSchedule and MESSAGE=. Check from the message why it is blocked even though resources remain.

In step 1, case-ok came up on the same node, but this Pod cannot go there. The only difference is one toleration. The message in this step differs from the previous two steps — a taint is one of the few spots where the scheduler tells you the cause precisely. You can see the taints on a node with kubectl get node lab-node-0 -o jsonpath='{.spec.taints}'.

A Pod with no candidate node at all

In /root/gputri/cases/case-label.yaml, write the Pod case-label — set the nodeSelector to nvidia.com/gpu.product: NVIDIA-A100-SXM4-40GB (no node has this label), and leave the toleration and the nvidia.com/gpu: 1 request as they are. After applying it, write two lines in /root/gputri/out/05-label.txt — for MATCHING_NODES=, the number of nodes that actually have that label, and for MESSAGE=, the condition message.

This is what happens when a manifest that chose a label attached by GPU Feature Discovery as it is meets a cluster that does not have that label. For this Pod, there is no taint or resource to weigh — because there are 0 candidate nodes. That is why nodeSelector comes first in the order of judgment. You can count the nodes that have the label with kubectl get nodes -l <키> --no-headers | wc -l (the placeholder is the key).

The advertisement is fine, but there are no slots

In /root/gputri/cases/case-exhaust.yaml, write the Pod case-exhaust — pin it to lab-node-0, put in the toleration, and require nvidia.com/gpu: 2. That node advertises 2 cards, but case-ok from step 1 already holds one. After applying it, write three lines in /root/gputri/out/06-exhaust.txt — ALLOCATABLE= (the amount that node advertised), ALLOCATED= (the sum of the GPU requests of the Pods placed on that node), and REQUESTED=2.

Only this judgment does not end with a single node object. You must gather the Pods placed on that node and add up their requests — kubectl get pods -A --field-selector spec.nodeName=lab-node-0 -o json is the starting point. Pods that have ended (Succeeded and Failed) do not hold slots, so they must be excluded from the sum. The message is again the same as in steps 2 and 3 — this is the third time you see the same sentence.

Two cases in which the Pod is not created at all

First, in /root/gputri/cases/case-norequest.yaml, write the Pod case-norequest — pin it to lab-node-0, put in the toleration, but do not put in a resource request at all. When applied, it comes up. Next, in /root/gputri/cases/case-rtc.yaml, write a Pod case-rtc carrying runtimeClassName: nvidia and apply it, and save the output, including standard error, to /root/gputri/out/07-runtimeclass.txt (do not create the RuntimeClass). Finally, create the namespace gpu-quota, write a ResourceQuota gpu-quota in /root/gputri/cases/quota.yaml that limits requests.nvidia.com/gpu to "1", then apply the Pod case-quota (requiring 3 GPUs) in /root/gputri/cases/case-quota.yaml and save the output to /root/gputri/out/07-quota.txt. Lastly, write three lines in /root/gputri/out/07-symptoms.txt — NOREQUEST=, RUNTIMECLASS=, and QUOTA=. On the latter two lines, write whether a Pod object was created as created or absent.

The three in this step have a different symptom strand from the first five. The first is a Pod that comes up fine but has no device attached, and the latter two are cases where the Pod object is not created at all. If it is rejected at admission, there is no trace in the cluster and the error remains only on the terminal of the person who applied it — so the habit of putting the output into a file with 2>&1 is decisive in an investigation. That a Pod does not exist can be confirmed by the exit code of kubectl get pod <이름> -n <ns> (the placeholder is the name).

A classifier that answers the cause in one word by looking only at objects

Create /root/gputri/bin/triage.sh. When called as bash triage.sh <네임스페이스> <파드> (the placeholders are the namespace and the Pod), it prints only one word to standard output and ends. There are seven answers — no-request, ok, label-mismatch, taint, no-gpu-node, allocatable-zero, and exhausted. If the Pod does not exist, it prints a message to standard error and ends with 1. The judgment is made only from the objects given by kubectl get -o json, and the order is no request → already placed → nodeSelector → taint → capacity → allocatable → remaining slots. After creating it, run it on all seven Pods (case-ok, case-nogpunode, case-alloc0, case-taint, case-label, case-exhaust, and case-norequest) and write seven lines in /root/gputri/out/triage.txt in the form <파드이름> <원인> (the placeholders are the Pod name and the cause).

You must not decide the answer by the Pod name — it is useful only if it is right when run on a Pod you see for the first time. There are three inputs you need: that Pod, all the nodes, and the Pods of all namespaces (for slot computation). The toleration judgment must handle both the operator: Exists and Equal cases, and if effect is empty, it tolerates all effects. For extended resources, looking only at limits is enough — because of the rule that requests and limits must be equal. Do not struggle to parse JSON in the shell; if you pass it to python3 it gets much shorter.