TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

The linter was quiet, but the buckets were backwards

Continue in TT Lab

Goal

You diagnose and fix exporter output that violates the conventions with the promtool linter, recount the histogram exposition from the raw data, load it into a TSDB with OpenMetrics and check the number of time series, and then build an exporter with only the standard library and actually scrape it.

Why it matters

Prometheus trusts the text an exporter puts out as it is. If a counter lacks _total or units are mixed in milliseconds, dashboards and rules read it with different meanings, and if buckets are not cumulative, quantiles are quietly wrong. You also have no feel for how many time series one label and a few buckets multiply into until you load them. Telling apart what the linter catches and what it cannot is the starting point of an instrumentation review.

Prepared environment

python3 /opt/fixtures/pca_exposition_lab.py init puts legacy.prom (old exporter output that violates the conventions) and payload-sizes.csv (40 request sizes) in /root/pca-exposition/. It actually runs check metrics (lint) and tsdb create-blocks-from openmetrics (block creation) with promtool 3.0.1 from the lab-k8s image. It does not start a Prometheus server. The grader briefly starts the exporters of steps 6 and 7 on a free port and then shuts them down, so the processes you started have nothing to do with grading.

Steps

  1. After creating the material with python3 /opt/fixtures/pca_exposition_lab.py init, run promtool check metrics < /root/pca-exposition/legacy.prom. In /root/pca-exposition/01-lint.txt, write problems= (the number of lines flagged), metrics= (the flagged metric names, comma-separated, without duplicates), and missed_by_lint= (the name of the series that the lint is quiet about but is wrong because its buckets are not cumulative).
  2. Create /root/pca-exposition/fixed.prom and move the five series of legacy.prom into it. http_requests becomes the counter http_requests_total (keep both samples), request_latency_ms 212 becomes the gauge request_latency_seconds 0.212 (add HELP), queueDepth becomes queue_depth, cache_hits_total becomes the gauge cache_entries, and job_duration_count becomes the gauge batch_jobs_running. Do not move payload_bytes. promtool check metrics must flag nothing.
  3. Using the request sizes in /root/pca-exposition/payload-sizes.csv, build the histogram http_request_size_bytes and append it to the end of fixed.prom. The boundaries are 100, 1000, 10000, and +Inf, and also write _sum and _count. Each bucket is the cumulative count of requests at or below (inclusive of) that boundary. The lint must stay clean.
  4. Move the same samples as fixed.prom into the OpenMetrics format /root/pca-exposition/scrape.om. Attach a timestamp (in seconds, e.g. 1700000000) to the end of every sample, and the last line is # EOF. After loading with promtool tsdb create-blocks-from openmetrics /root/pca-exposition/scrape.om /root/pca-exposition/tsdb, write series= (NUM SERIES) and samples= (NUM SAMPLES) in /root/pca-exposition/04-import.txt.
  5. In /root/pca-exposition/cardinality.om, write the http_request_size_bytes histogram with three route label values, /checkout, /search, and /upload, using 12 finite boundaries and +Inf (including _sum and _count for each route, timestamps, and # EOF). Before loading, calculate the number of time series and write it as predicted_series= in /root/pca-exposition/05-cardinality.txt, and write the load result as imported_series=.
  6. Write /root/pca-exposition/exporter.py with only the standard library. It listens on 127.0.0.1 at the environment variable PORT (9464 if not set), and for each GET /work?d=초 (the placeholder is the seconds), it increments the counter demo_jobs_processed_total by 1 and records d in the histogram demo_job_duration_seconds (boundaries 0.1, 0.5, 1, 5). GET /metrics returns an exposition format with HELP and TYPE, with Content-Type: text/plain; version=0.0.4. The grader starts it fresh on a free port, sends d=0.05, 0.5, 2, and 7, scrapes it, checks the lint, the counter increase, the cumulative buckets, and _sum, and then shuts it down.
  7. Attach a queue label to the counter in exporter.py. Use the name from GET /work?d=초&queue=이름 (the placeholders are the seconds and the name), and if it is absent, use default. Escape backslashes, double quotes, and newlines in label values according to the exposition format rules. The grader sends the names exports, say "hi", c:\tmp, and one containing a newline, and checks that promtool accepts the exposition and that the four names are each read back once with their original values.

Notes

The six lines promtool rejected

After creating the material with python3 /opt/fixtures/pca_exposition_lab.py init, run promtool check metrics < /root/pca-exposition/legacy.prom. In /root/pca-exposition/01-lint.txt, write problems= (the number of lines flagged), metrics= (the flagged metric names, comma-separated, without duplicates), and missed_by_lint= (the name of the series that the lint is quiet about but is wrong because its buckets are not cumulative).

The linter looks at naming, HELP, and suffix conventions. It does not check whether histogram buckets are cumulative, so read the _bucket values in le order yourself.

Bringing names and units into line with the conventions

Create /root/pca-exposition/fixed.prom and move the five series of legacy.prom into it. http_requests becomes the counter http_requests_total (keep both samples), request_latency_ms 212 becomes the gauge request_latency_seconds 0.212 (add HELP), queueDepth becomes queue_depth, cache_hits_total becomes the gauge cache_entries, and job_duration_count becomes the gauge batch_jobs_running. Do not move payload_bytes. promtool check metrics must flag nothing.

Only counters end in _total. _count, _sum, and _bucket are suffixes used by histograms and summaries. Change units to the base unit (seconds, bytes) and convert the values along with them.

Recounting cumulative buckets from the CSV

Using the request sizes in /root/pca-exposition/payload-sizes.csv, build the histogram http_request_size_bytes and append it to the end of fixed.prom. The boundaries are 100, 1000, 10000, and +Inf, and also write _sum and _count. Each bucket is the cumulative count of requests at or below (inclusive of) that boundary. The lint must stay clean.

The payload_bytes in legacy wrote per-interval counts and was not cumulative. le="+Inf" is always equal to _count.

Loading with OpenMetrics and counting time series

Move the same samples as fixed.prom into the OpenMetrics format /root/pca-exposition/scrape.om. Attach a timestamp (in seconds, e.g. 1700000000) to the end of every sample, and the last line is # EOF. After loading with promtool tsdb create-blocks-from openmetrics /root/pca-exposition/scrape.om /root/pca-exposition/tsdb, write series= (NUM SERIES) and samples= (NUM SAMPLES) in /root/pca-exposition/04-import.txt.

OpenMetrics rejects loading if there is no # EOF or no timestamp. One histogram becomes a time series for each bucket, and _sum and _count each become one too.

The price of one label and 12 buckets

In /root/pca-exposition/cardinality.om, write the http_request_size_bytes histogram with three route label values, /checkout, /search, and /upload, using 12 finite boundaries and +Inf (including _sum and _count for each route, timestamps, and # EOF). Before loading, calculate the number of time series and write it as predicted_series= in /root/pca-exposition/05-cardinality.txt, and write the load result as imported_series=.

The number of time series is the product of label value combinations. One histogram produces (number of boundaries + 1) buckets plus _sum and _count.

Scraping an exporter you built yourself

Write /root/pca-exposition/exporter.py with only the standard library. It listens on 127.0.0.1 at the environment variable PORT (9464 if not set), and for each GET /work?d=초 (the placeholder is the seconds), it increments the counter demo_jobs_processed_total by 1 and records d in the histogram demo_job_duration_seconds (boundaries 0.1, 0.5, 1, 5). GET /metrics returns an exposition format with HELP and TYPE, with Content-Type: text/plain; version=0.0.4. The grader starts it fresh on a free port, sends d=0.05, 0.5, 2, and 7, scrapes it, checks the lint, the counter increase, the cumulative buckets, and _sum, and then shuts it down.

Count buckets cumulatively at or below the boundary. A single request is added to every boundary larger than itself and to +Inf. The grader gives you PORT, so do not fix the port.

A queue name arrived with quotation marks in it

Attach a queue label to the counter in exporter.py. Use the name from GET /work?d=초&queue=이름 (the placeholders are the seconds and the name), and if it is absent, use default. Escape backslashes, double quotes, and newlines in label values according to the exposition format rules. The grader sends the names exports, say "hi", c:\tmp, and one containing a newline, and checks that promtool accepts the exposition and that the four names are each read back once with their original values.

In an exposition format label value, only three characters are escaped. The order matters — if you do not replace the backslash first, the backslashes you newly inserted get replaced again.