TT Lab
Get started
Learn Learning paths Courses

GPU Operator and Time-Slicing

DCGM Exporter metrics, and what those metrics do not tell you

Continue in TT Lab

In one line

The raw material of GPU observability is the plain text that DCGM Exporter emits at /metrics, and the most important thing in that text is not the values but the labels. The labels connect a metric to a card, a host, and a Pod, and only with that connection can you prove with numbers the sentence "the allocation is all out but nobody is using it".

Why this was needed

The GPU is the most expensive resource in an organization, yet it is also the least visible. For cpu and memory, kubelet counts them by itself and they show immediately with kubectl top, but for the GPU, Kubernetes does not know the utilization. All Kubernetes knows is "one card was assigned to this Pod"; whether that one card is running at 100% or idle is none of its concern.

So a very common state arises. The node's nvidia.com/gpu allocatable is all used up and new jobs stay Pending, yet the actual utilization of the cards is a single digit. From the scheduler's point of view this cluster is a full cluster, and from the point of view of electricity and depreciation it is a cluster that is nearly idle. Since the two numbers live in different systems and nobody looks at them side by side, this state lasts for months.

The allocation rate comes from the Kubernetes API and the utilization comes from DCGM. What this module aims to do is put those two on the same screen.

How it works

DCGM Exporter comes up as a DaemonSet on each node, reads values from the cards with NVIDIA's DCGM library, and emits them in the Prometheus exposition format. The actual output in the official documentation looks like this.

# HELP DCGM_FI_DEV_GPU_TEMP GPU temperature (in C).
# TYPE DCGM_FI_DEV_GPU_TEMP gauge
DCGM_FI_DEV_GPU_TEMP{gpu="0",UUID="GPU-34319582-d595-d1c7-d1d2-179bcfa61660",device="nvidia0",Hostname="ub20-a100-k8s"} 61
DCGM_FI_DEV_FB_FREE{gpu="0",UUID="GPU-34319582-d595-d1c7-d1d2-179bcfa61660",device="nvidia0",Hostname="ub20-a100-k8s"} 35690
DCGM_FI_DEV_FB_USED{gpu="0",UUID="GPU-34319582-d595-d1c7-d1d2-179bcfa61660",device="nvidia0",Hostname="ub20-a100-k8s"} 4845
DCGM_FI_DEV_XID_ERRORS{gpu="0",UUID="GPU-34319582-d595-d1c7-d1d2-179bcfa61660",device="nvidia0",Hostname="ub20-a100-k8s"} 0
DCGM_FI_PROF_GR_ENGINE_ACTIVE{gpu="0",UUID="GPU-34319582-d595-d1c7-d1d2-179bcfa61660",device="nvidia0",Hostname="ub20-a100-k8s"} 0.995630

There are three things to hold on to here.

First, the names are the DCGM field names as they are. Those starting with DCGM_FI_DEV_ are device-level values (temperature, memory, XID errors, clocks), and those starting with DCGM_FI_PROF_ are profiling values (graphics engine active ratio, tensor pipe active ratio, DRAM active ratio). DCGM_FI_DEV_GPU_UTIL, which is commonly used to see utilization, is a coarse number closer to "was a kernel running at the sampling instant", while DCGM_FI_PROF_GR_ENGINE_ACTIVE gives, as a ratio from 0 to 1, how active the engine actually was during that interval. A job that leaves a slot idle for a long time looks busy by the former number and idle by the latter. Which number you use in capacity planning changes the conclusion.

Second, memory comes out as two metrics. They are DCGM_FI_DEV_FB_USED and DCGM_FI_DEV_FB_FREE (in MiB). Rather than looking for a total memory metric separately, it is usual to get the total by adding the two. The reason these two matter is that a large share of GPU failures appear in the shape "utilization is 0 but memory is held". A dead process has not released its context, and that card stays alive yet returns to nobody.

Third, the labels are half of this text. gpu is the index within the node, UUID is the card's global identifier, device is the device node name, and Hostname is the node. If you turn on the exporter's Kubernetes mapping (the environment variable DCGM_EXPORTER_KUBERNETES, the flag -k), information about the Pod holding that card is attached as additional labels. Only with this mapping can you answer the question "who is holding this card", and without it the metrics talk only about cards and cannot talk about people.

If you turn on MIG, the same metrics come out at both the card level and the GPU instance level, and the instance lines get extra GPU_I_PROFILE and GPU_I_ID labels. So if you carelessly sum() the metrics in a MIG cluster, you add the same value twice. Clusters whose dashboard total is larger than the number of physical cards are nearly always this.

Collection path and alerts

For Prometheus to scrape the metrics, a target definition is needed. Where the Prometheus Operator is used, that definition is the ServiceMonitor custom resource. If you write which port of which service to scrape, at what interval, from which path, the Operator makes the Prometheus configuration for you. What matters is that this is an API object with a schema. If you write interval as 30 instead of 30s, the API server rejects it as a type mismatch. A typo that would have been silently ignored in a configuration file gets caught at apply time.

For what to alert on, do not alert because a value is large; choose things for which the human's job is decided.

Alert Basis metric Why a person must move
XID error occurred DCGM_FI_DEV_XID_ERRORS It is a hardware or driver event. You must pull the card or empty the node
Temperature threshold exceeded DCGM_FI_DEV_GPU_TEMP It is a cooling problem, and if neglected the clock drops and performance quietly decreases
Memory held with no owner DCGM_FI_DEV_FB_USED and the absence of a Pod label A dead process is holding the card. The node needs attention
Allocated but idle The gap between the allocation (API) and DCGM_FI_DEV_GPU_UTIL It is a cost problem. You must ask the job owner to give it back

The last row is the core of this module, and it is an alert that cannot be made from a single metric. One side must come from the Kubernetes API and the other from the exporter.

What it looks like in the field

First, the organization is surprised on the day the utilization dashboard is first turned on. Average utilization 15%, yet the queue is long. Usually the cause is not technology but habit. People leave notebooks up and go home, and a single job holds a whole card while doing data preprocessing on the CPU. Until shown the numbers, nobody knows they are doing that.

Second, if you turn on time-slicing, the metrics get confusing. Time-slicing only increases the number of resource names and does not divide the card. So the metrics of a node using five slots still come out as a single-card line. Even if you want to split utilization per Pod, there is no such value. It is common to waste time trying to build "utilization per slot" without knowing this. On time-slicing nodes, look at the card-level metrics and the Pod-level allocation separately, and do not interpret them by multiplying the two.

Third, label cardinality quietly becomes a problem. Each metric gets a UUID and a Pod name attached, so in a cluster with many short-lived Pods, time series are continually newly created. If you put the time series containing Pod names straight into long-term retention, the storage collapses first. A separation is needed, such as reducing long-term retention to the card level and keeping the Pod level short.

Fourth, the honest limit of this environment. The lab Pod has no GPU, no DCGM, and no Prometheus. So in the next lab you write exposition-format text by hand and use it as the material. It may look like imitation, but a large share of the real operations work is exactly this — reading a /metrics dump someone else gave you, parsing it, and cross-checking whether the names that rules reference actually come out. PromQL itself cannot be executed without Prometheus, so it is not executed, and you judge only as far as the structure of the rules file and the existence of the metric names it references.

Reference documents

What you will do in the next lab

You write DCGM Exporter's exposition-format text by hand. You write five metrics for four cards complete with labels, and build a tool that parses it into a per-card table. You make a separate version with the Kubernetes mapping turned on and find a card that holds memory but has no Pod label, and count the same cluster's allocation from the API and put it side by side with utilization. You write an alert rules file and build a checker that cross-checks whether the metric names the rules reference actually appear in the exposition. Finally, you define the ServiceMonitor CRD yourself and put it into the API server, and see how a wrongly written interval gets rejected.