TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Judging the Idle Card from Four Real GPUs

Continue in TT Lab

Goal

With three hours of real measurements from four GPUs, you judge "is this card idle" and leave the grounds for that verdict as a report. At the end, you derive time to first token from serving metrics in the vLLM format.

Why it matters

GPUs are expensive and do not increase easily, so a wrong capacity verdict directly leaks money and time. And if you look only at utilization, you are always wrong — if a card with utilization 0 holds 13.7 GiB of memory, that card is idle and full at the same time. On top of that, the attribution labels that tell you which Pod is using it can be entirely empty depending on the node configuration, so a "GPU whose user is unknown" really occurs. This lab is practice in judging those two things from metrics.

About the data

Steps

  1. Write a query that counts the number of GPUs DCGM reports in /root/pca-gpu/01-inventory.promql.
  2. Write on one line, in /root/pca-gpu/02-idle.txt, the node name of the GPU that has utilization 0 throughout the window and holds the most framebuffer.
  3. Write a query that gives that GPU's DCGM_FI_DEV_FB_USED in /root/pca-gpu/03-fbused.promql.
  4. Write on one line, in /root/pca-gpu/04-attribution.txt, the node name of the GPU whose exported_pod label is empty.
  5. Write a query that gives the average power (W) from the cumulative energy counter of the nuc1 GPU in /root/pca-gpu/05-energy.promql.
  6. Write a query that counts the number of GPUs whose utilization never exceeded 0 over the whole window in /root/pca-gpu/06-idle-window.promql.
  7. Write a query that gets the p95 of time to first token for the serving samples in /root/pca-gpu/07-ttft.promql.
  8. Write the two verdicts and their grounds, and which parts are measured and which are samples, in /root/pca-gpu/08-report.md in at least 300 characters.

Notes

Counting the GPUs DCGM reports

DCGM emits one series per GPU. Just count with count(). Throw the query with promq-at — if you throw it at the present time, it is outside the window and an empty result comes back.

Finding a GPU that holds memory but is idle throughout the window

If you judge by a single instantaneous value, even a busy card gets caught at a moment when it happens to be 0. With max_over_time(...[3h]), pick the GPUs whose maximum over the whole window is 0, and find the one with the largest FB_USED among them with topk. The answer is a node label value.

Measuring the framebuffer that GPU holds

The unit of DCGM_FI_DEV_FB_USED is MiB. Narrow to the node you found in the previous step. If you convert to bytes, it will not match the reference value — answer in the unit the metric gives.

Finding a GPU whose user cannot be known

dcgm-exporter tells you which Pod uses that GPU through the exported_namespace, exported_pod, and exported_container labels. But depending on the node configuration these labels can be empty. You pick an empty label with {exported_pod=""}.

Deriving average power from a cumulative energy counter

DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION is cumulative energy and its unit is mJ. If you differentiate with rate it is mW, and dividing by 1000 gives W. The node is nuc1. Compared with the POWER_USAGE gauge at the same time, the values should be almost the same if it is normal.

Counting idle GPUs over the whole window

If you count by instantaneous value, even a busy GPU gets caught at a moment when it happens to be 0. Get the maximum over the whole window with max_over_time, then filter with == 0 and count.

Getting the p95 of time to first token from serving metrics

The series that start with vllm: are not real measurements but synthetic samples imitating the format (they carry the label source="synthetic"). Since it is a histogram, apply rate to the buckets, aggregate by le, and then wrap with histogram_quantile. If you drop le when aggregating, the value becomes meaningless.

Writing the verdicts and grounds, together with what is measured

The report must include the two verdicts (the node of the GPU that only holds memory and is idle, and the node of the GPU whose attribution cannot be known) and the names of the metrics you used as grounds. Also write that the vllm: series are synthetic samples, not real measurements. It is at least 300 characters.