PCA — Prometheus Certified Associate
Judging the Idle Card from Four Real GPUs
Goal
With three hours of real measurements from four GPUs, you judge "is this card idle" and leave the grounds for that verdict as a report. At the end, you derive time to first token from serving metrics in the vLLM format.
Why it matters
GPUs are expensive and do not increase easily, so a wrong capacity verdict directly leaks money and time. And if you look only at utilization, you are always wrong — if a card with utilization 0 holds 13.7 GiB of memory, that card is idle and full at the same time. On top of that, the attribution labels that tell you which Pod is using it can be entirely empty depending on the node configuration, so a "GPU whose user is unknown" really occurs. This lab is practice in judging those two things from metrics.
About the data
- The GPU, container, and node metrics are real measurements from the 7-node cluster that runs this site.
- Only the series that start with
vllm:are not real measurements. The vLLM deployment on this cluster is currently off and has no metrics. They are synthetic samples that only imitate the format, and every series has the labelsource="synthetic". - The window is a fixed 3 hours. Check it with
lab-realdataand query withpromq-at. - The original exposition text is in
/opt/lab/realdata/expo/, and the source description is inREADME.mdin the same directory.
Steps
- Write a query that counts the number of GPUs DCGM reports in
/root/pca-gpu/01-inventory.promql. - Write on one line, in
/root/pca-gpu/02-idle.txt, the node name of the GPU that has utilization 0 throughout the window and holds the most framebuffer. - Write a query that gives that GPU's
DCGM_FI_DEV_FB_USEDin/root/pca-gpu/03-fbused.promql. - Write on one line, in
/root/pca-gpu/04-attribution.txt, the node name of the GPU whoseexported_podlabel is empty. - Write a query that gives the average power (W) from the cumulative energy counter of the
nuc1GPU in/root/pca-gpu/05-energy.promql. - Write a query that counts the number of GPUs whose utilization never exceeded 0 over the whole window in
/root/pca-gpu/06-idle-window.promql. - Write a query that gets the p95 of time to first token for the serving samples in
/root/pca-gpu/07-ttft.promql. - Write the two verdicts and their grounds, and which parts are measured and which are samples, in
/root/pca-gpu/08-report.mdin at least 300 characters.
Notes
- If you throw a query like
promq-at 'DCGM_FI_DEV_GPU_UTIL', the labels are shown as they are. The node name is in thenodelabel. - Common mistake 1: judging with an instantaneous value without
max_over_time. Even a busy card has moments when it happens to be 0. - Common mistake 2: converting
FB_USEDto bytes. The unit of this metric is MiB.
Counting the GPUs DCGM reports
DCGM emits one series per GPU. Just count with count(). Throw the query with promq-at — if you throw it at the present time, it is outside the window and an empty result comes back.
Finding a GPU that holds memory but is idle throughout the window
If you judge by a single instantaneous value, even a busy card gets caught at a moment when it happens to be 0. With max_over_time(...[3h]), pick the GPUs whose maximum over the whole window is 0, and find the one with the largest FB_USED among them with topk. The answer is a node label value.
Measuring the framebuffer that GPU holds
The unit of DCGM_FI_DEV_FB_USED is MiB. Narrow to the node you found in the previous step. If you convert to bytes, it will not match the reference value — answer in the unit the metric gives.
Finding a GPU whose user cannot be known
dcgm-exporter tells you which Pod uses that GPU through the exported_namespace, exported_pod, and exported_container labels. But depending on the node configuration these labels can be empty. You pick an empty label with {exported_pod=""}.
Deriving average power from a cumulative energy counter
DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION is cumulative energy and its unit is mJ. If you differentiate with rate it is mW, and dividing by 1000 gives W. The node is nuc1. Compared with the POWER_USAGE gauge at the same time, the values should be almost the same if it is normal.
Counting idle GPUs over the whole window
If you count by instantaneous value, even a busy GPU gets caught at a moment when it happens to be 0. Get the maximum over the whole window with max_over_time, then filter with == 0 and count.
Getting the p95 of time to first token from serving metrics
The series that start with vllm: are not real measurements but synthetic samples imitating the format (they carry the label source="synthetic"). Since it is a histogram, apply rate to the buckets, aggregate by le, and then wrap with histogram_quantile. If you drop le when aggregating, the value becomes meaningless.
Writing the verdicts and grounds, together with what is measured
The report must include the two verdicts (the node of the GPU that only holds memory and is idle, and the node of the GPU whose attribution cannot be known) and the names of the metrics you used as grounds. Also write that the vllm: series are synthetic samples, not real measurements. It is at least 300 characters.