GPU metrics — write the exposition, parse it, wire it to alerts
Goal
You write DCGM Exporter's exposition format by hand, parse it into a per-card table, put the allocation rate and the utilization side by side, and go as far as making alert rules and a scrape target object.
Why it matters
The GPU is the most expensive resource in an organization, yet Kubernetes does not know its utilization. All it knows is "one card was assigned to this Pod". So a cluster that looks full to the scheduler can, from the electricity bill's point of view, stay nearly idle for months. To turn this gap into numbers, you must put the allocation from the Kubernetes API and the utilization from DCGM on the same screen, and to do that you must first know the actual shape of the metrics. What matters in metrics is not the values but the labels — labels connect a metric to a card, a node, and a Pod, and failures such as "a card holding memory with no owner" come to light at the spot where that connection is broken.
Steps
- Create
/root/gpumet/metrics,/root/gpumet/bin,/root/gpumet/out,/root/gpumet/rules, and/root/gpumet/k8s, and in/root/gpumet/metrics/dcgm.promwrite the/metricsoutput of a node with 4 A100s plugged in. There are five metrics —DCGM_FI_DEV_GPU_UTIL,DCGM_FI_DEV_FB_USED,DCGM_FI_DEV_FB_FREE,DCGM_FI_DEV_GPU_TEMP, andDCGM_FI_DEV_XID_ERRORS. For each metric, put one# HELPline and one# TYPE <이름> gaugeline (the placeholder is the name) in front, and follow them with the samples for the four cards. The sample labels are four —gpu(0 to 3),UUID(GPU-a1b2c3d4-0000-0000-0000-00000000000<번호>, where the placeholder is the number),device(nvidia<번호>, where the placeholder is the number), andHostname(lab-node-0). The values are: for gpu 0, utilization 97 · FB_USED 38000 · FB_FREE 2760 · temperature 71 · XID 0; for gpu 1, 2 · 9000 · 31760 · 41 · 2; for gpu 2, 0 · 0 · 40760 · 33 · 0; and for gpu 3, 1 · 21000 · 19760 · 40 · 0. - Create
/root/gpumet/bin/parse-metrics.sh <노출파일>(the placeholder is the exposition file). For each card it prints one line with five columns separated by spaces,<gpu> <util> <fb_used> <fb_total> <mem_pct>.fb_totalis the sum of FB_USED and FB_FREE, andmem_pctis the truncated integer offb_used * 100 / fb_total. The lines must be in ascending order of gpu number, and it must read only the file it receives as an argument (do not bake a path into it). After creating it, run it on the step 1 file and save the output to/root/gpumet/out/per-gpu.txt. - In
/root/gpumet/metrics/dcgm-k8s.prom, write the version with the Kubernetes mapping turned on — the same five metrics and four cards as step 1, with three more labels added only to gpu 0 and gpu 3 (namespace="gpu-metrics"; for gpu 0,pod="train-a"andcontainer="trainer"; for gpu 3,pod="infer-b"andcontainer="server"). gpu 1 and gpu 2 have no Pod labels. And create/root/gpumet/bin/find-orphan.sh <노출파일>(the placeholder is the exposition file) — it prints, one per line in ascending order, the gpu numbers of cards that have no Pod label but whose FB_USED exceeds 1024. Run it on this new file, not the step 1 file, and save the output to/root/gpumet/out/orphan.txt. - In both
status.capacityandstatus.allocatableoflab-node-0, putnvidia.com/gpuas"4", create the namespacegpu-metrics, and in/root/gpumet/k8s/workloads.yamlwrite three Pods (train-a,infer-b, andidle-c) — all pinned tolab-node-0, each requiringnvidia.com/gpu: 1. When all three are up after applying, write five lines in/root/gpumet/out/gap.txt—PHYSICAL=(the number of cards the node advertised),ALLOCATED=(the number of cards that went out as the Pods requested),ALLOC_PCT=(the ratio of the two, truncated integer),MEAN_UTIL=(the average utilization of the four cards in the step 3 exposition file, truncated integer), andIDLE_BUT_ALLOCATED=(the number of cards that have a Pod label but utilization below 10). - In
/root/gpumet/rules/gpu-alerts.yaml, write a Prometheus rules file. There isgroupsat the top level, thenameof one group isgpu, and inside it are exactly three rules —GpuXidError(an expression usingDCGM_FI_DEV_XID_ERRORS,for: 0m,severity: critical),GpuTempHigh(an expression usingDCGM_FI_DEV_GPU_TEMP,for: 10m,severity: warning), andGpuAllocatedButIdle(an expression usingDCGM_FI_DEV_GPU_UTIL,for: 2h,severity: info). Give each rule all five ofalert,expr,for,labels.severity, andannotations.summary. Writesummaryas one sentence that contains "what to do". - Create
/root/gpumet/bin/check-rules.sh <규칙파일> <노출파일>(the placeholders are the rules file and the exposition file). From everyexprin the rules file it extracts the metric names starting withDCGM_FI_, and if there is not a single sample of that name in the exposition file, it prints one lineMISSING=<이름>(the placeholder is the name) for each and ends with1. If all are present, it printsOK=<노출에 있는 지표 수>(the placeholder is the number of metrics in the exposition) and ends with0. After creating it, run it on the step 5 rules and the step 1 exposition file and save the output to/root/gpumet/out/rulecheck.txt. - In
/root/gpumet/k8s/servicemonitor-crd.yaml, write and apply theservicemonitors.monitoring.coreos.comCRD — groupmonitoring.coreos.com, versionv1, namespace scope, kindServiceMonitor. The schema isspec.selector.matchLabels(a string map) andspec.endpoints(an array,minItems: 1, where each item has a required stringport, a stringpath, andintervalas a string with the pattern^[0-9]+(ms|s|m|h)$), and inspec, bothselectorandendpointsare required. Next, in/root/gpumet/k8s/servicemonitor.yaml, write and applynvidia-dcgm-exporterin thegpu-metricsnamespace (selectorapp: nvidia-dcgm-exporter, one endpoint withport: gpu-metrics,path: /metrics,interval: 15s). Finally, in/root/gpumet/k8s/servicemonitor-bad.yaml, write the same thing with the namedcgm-bad-intervalbut withintervalas an unquoted30, apply it, and save the rejection output, including standard error, to/root/gpumet/out/07-reject.txt. - Create
/root/gpumet/bin/gpu-report.sh <노출파일> <규칙파일>(the placeholders are the exposition file and the rules file). It prints six lines in order —GPUS=(the number of cards that appear in the exposition),MEAN_UTIL=(average utilization, truncated),MAX_TEMP=(the highest temperature),ORPHAN_GPUS=(the number of cards with no Pod label whose FB_USED exceeds 1024),XID_GPUS=(the number of cards whose XID errors are greater than 0), andALERT_RULES=(the total number of rules in the rules file). All the numbers must be computed from the two files it receives as arguments. Run it on the step 3 exposition file and the step 5 rules file and save the output to/root/gpumet/out/report.txt.
Notes
- This Pod has neither a GPU, nor DCGM, nor Prometheus. So you make the exposition text by hand in step 1 and the later steps use it as material. The metric names and label names are exactly the actual output in the official documentation.
- PromQL is not executed. Alert rules are judged only as far as the structure and the existence of the metric names they reference.
- The shape of one exposition line:
이름{라벨="값",…} 숫자(name, labels, and number). There are only two kinds of comments,# HELPand# TYPE. - There is no total memory metric — you get it by adding FB_USED and FB_FREE.
- Common mistake: if you round the ratios, the grader catches it. Every ratio in this lab is a truncated integer.
- Common mistake: if you bake the numbers or paths of the step 1 file into a script, it fails when the grader runs it on a different file.
Write DCGM Exporter's exposition format by hand
Create /root/gpumet/metrics, /root/gpumet/bin, /root/gpumet/out, /root/gpumet/rules, and /root/gpumet/k8s, and in /root/gpumet/metrics/dcgm.prom write the /metrics output of a node with 4 A100s plugged in. There are five metrics — DCGM_FI_DEV_GPU_UTIL, DCGM_FI_DEV_FB_USED, DCGM_FI_DEV_FB_FREE, DCGM_FI_DEV_GPU_TEMP, and DCGM_FI_DEV_XID_ERRORS. For each metric, put one # HELP line and one # TYPE <이름> gauge line (the placeholder is the name) in front, and follow them with the samples for the four cards. The sample labels are four — gpu (0 to 3), UUID (GPU-a1b2c3d4-0000-0000-0000-00000000000<번호>, where the placeholder is the number), device (nvidia<번호>, where the placeholder is the number), and Hostname (lab-node-0). The values are: for gpu 0, utilization 97 · FB_USED 38000 · FB_FREE 2760 · temperature 71 · XID 0; for gpu 1, 2 · 9000 · 31760 · 41 · 2; for gpu 2, 0 · 0 · 40760 · 33 · 0; and for gpu 3, 1 · 21000 · 19760 · 40 · 0.
The Prometheus exposition format is one line of 이름{라벨="값",…} 숫자 (name, labels, and number). The two comment lines (# HELP and # TYPE) appear only once per metric, and after them many sample lines differing only in labels follow. Note that there is no separate total memory metric — you must add FB_USED and FB_FREE to get the total, and in this table the sum is the same for all four cards. Instead of typing 20 lines by hand, you may generate them with a shell loop.
Turn the exposition text into a per-card table
Create /root/gpumet/bin/parse-metrics.sh <노출파일> (the placeholder is the exposition file). For each card it prints one line with five columns separated by spaces, <gpu> <util> <fb_used> <fb_total> <mem_pct>. fb_total is the sum of FB_USED and FB_FREE, and mem_pct is the truncated integer of fb_used * 100 / fb_total. The lines must be in ascending order of gpu number, and it must read only the file it receives as an argument (do not bake a path into it). After creating it, run it on the step 1 file and save the output to /root/gpumet/out/per-gpu.txt.
Extract gpu="…" from the label string to tell the cards apart, and gather values by metric name. One regular expression can split 이름{라벨} 값 (name, labels, and value) into three pieces. A truncated integer comes straight out of Python's // — if you round, the grader catches it. The grader also runs this script on a file with different values. So you must not bake the numbers of the step 1 file in as the answer.
Attach Pod labels and find memory held with no owner
In /root/gpumet/metrics/dcgm-k8s.prom, write the version with the Kubernetes mapping turned on — the same five metrics and four cards as step 1, with three more labels added only to gpu 0 and gpu 3 (namespace="gpu-metrics"; for gpu 0, pod="train-a" and container="trainer"; for gpu 3, pod="infer-b" and container="server"). gpu 1 and gpu 2 have no Pod labels. And create /root/gpumet/bin/find-orphan.sh <노출파일> (the placeholder is the exposition file) — it prints, one per line in ascending order, the gpu numbers of cards that have no Pod label but whose FB_USED exceeds 1024. Run it on this new file, not the step 1 file, and save the output to /root/gpumet/out/orphan.txt.
If a dead process does not release its context, memory is held but no Pod holds that card. Kubernetes sees that card as empty and sends a new Pod, and that Pod fails with a memory shortage. There are two signals to use for the judgment — the presence of the pod label and the value of FB_USED. Take care that lines with the label and lines without do not get mixed on the same card. In this lab it is consistent per card. The grader also runs it on a file with different values.
Put the allocation rate and the utilization on the same screen
In both status.capacity and status.allocatable of lab-node-0, put nvidia.com/gpu as "4", create the namespace gpu-metrics, and in /root/gpumet/k8s/workloads.yaml write three Pods (train-a, infer-b, and idle-c) — all pinned to lab-node-0, each requiring nvidia.com/gpu: 1. When all three are up after applying, write five lines in /root/gpumet/out/gap.txt — PHYSICAL= (the number of cards the node advertised), ALLOCATED= (the number of cards that went out as the Pods requested), ALLOC_PCT= (the ratio of the two, truncated integer), MEAN_UTIL= (the average utilization of the four cards in the step 3 exposition file, truncated integer), and IDLE_BUT_ALLOCATED= (the number of cards that have a Pod label but utilization below 10).
The point of this step is that the two numbers come from different systems. The first three lines come from the Kubernetes API and the last two from the exposition file. Neither alone can say "it is idling expensively". Count the allocation by adding up the limits of kubectl get pods -n <ns> -o json (the placeholder is the namespace). Do not round the average; compute it by truncation.
Alert only on what a person must move for
In /root/gpumet/rules/gpu-alerts.yaml, write a Prometheus rules file. There is groups at the top level, the name of one group is gpu, and inside it are exactly three rules — GpuXidError (an expression using DCGM_FI_DEV_XID_ERRORS, for: 0m, severity: critical), GpuTempHigh (an expression using DCGM_FI_DEV_GPU_TEMP, for: 10m, severity: warning), and GpuAllocatedButIdle (an expression using DCGM_FI_DEV_GPU_UTIL, for: 2h, severity: info). Give each rule all five of alert, expr, for, labels.severity, and annotations.summary. Write summary as one sentence that contains "what to do".
You set an alert not because a value is large but when the person's job is decided. XID is a hardware event so you must empty the node, temperature is a cooling problem, and an idle allocation is a cost problem so you must contact the job owner. That is also why the for of the three rules differs — XID immediately after a single occurrence, utilization after watching for a long while. This file is not applied to the cluster. It is material for the checker in the next step to read.
Cross-check that the metrics the rules call actually appear
Create /root/gpumet/bin/check-rules.sh <규칙파일> <노출파일> (the placeholders are the rules file and the exposition file). From every expr in the rules file it extracts the metric names starting with DCGM_FI_, and if there is not a single sample of that name in the exposition file, it prints one line MISSING=<이름> (the placeholder is the name) for each and ends with 1. If all are present, it prints OK=<노출에 있는 지표 수> (the placeholder is the number of metrics in the exposition) and ends with 0. After creating it, run it on the step 5 rules and the step 1 exposition file and save the output to /root/gpumet/out/rulecheck.txt.
If you get one character of a metric name wrong in an alert rule, that alert never fires. That is because Prometheus does not see a nonexistent metric as an error but as an empty result. So this cross-check is worth hanging before deployment. Plausible typos such as DCGM_FI_DEV_GPU_UTILIZATION really happen often. You can extract the names in the exposition file from lines that start with 이름{ (the name followed by a brace), and must not count names in comment lines. The grader also runs it on a deliberately wrong rules file.
Make the scrape target an object with a schema
In /root/gpumet/k8s/servicemonitor-crd.yaml, write and apply the servicemonitors.monitoring.coreos.com CRD — group monitoring.coreos.com, version v1, namespace scope, kind ServiceMonitor. The schema is spec.selector.matchLabels (a string map) and spec.endpoints (an array, minItems: 1, where each item has a required string port, a string path, and interval as a string with the pattern ^[0-9]+(ms|s|m|h)$), and in spec, both selector and endpoints are required. Next, in /root/gpumet/k8s/servicemonitor.yaml, write and apply nvidia-dcgm-exporter in the gpu-metrics namespace (selector app: nvidia-dcgm-exporter, one endpoint with port: gpu-metrics, path: /metrics, interval: 15s). Finally, in /root/gpumet/k8s/servicemonitor-bad.yaml, write the same thing with the name dcgm-bad-interval but with interval as an unquoted 30, apply it, and save the rejection output, including standard error, to /root/gpumet/out/07-reject.txt.
This is the benefit of handling it as a CRD — a typo that would have been silently ignored in a configuration file gets caught at apply time. In a structural schema, minItems is the lower bound of the array length and pattern is a regular expression constraint on a string. In YAML, an unquoted 30 is read as an integer and "30s" as a string — that difference is the whole of this step. After applying the CRD, it takes a round trip or two for the API server to accept the new type.
Compute a six-line GPU status report
Create /root/gpumet/bin/gpu-report.sh <노출파일> <규칙파일> (the placeholders are the exposition file and the rules file). It prints six lines in order — GPUS= (the number of cards that appear in the exposition), MEAN_UTIL= (average utilization, truncated), MAX_TEMP= (the highest temperature), ORPHAN_GPUS= (the number of cards with no Pod label whose FB_USED exceeds 1024), XID_GPUS= (the number of cards whose XID errors are greater than 0), and ALERT_RULES= (the total number of rules in the rules file). All the numbers must be computed from the two files it receives as arguments. Run it on the step 3 exposition file and the step 5 rules file and save the output to /root/gpumet/out/report.txt.
The two scripts you made in earlier steps already do half of it — here you gather those computations into one tool. The number of cards is the count of distinct values of the gpu label, not the number of sample lines. Do not round the average; give it by truncation. The grader also runs this script on a different exposition file and a different rules file — so if you bake in numbers, it fails. Cards with Pod labels and cards without are mixed, so gather and judge the presence of labels per card.