PCA — Prometheus Certified Associate
The linter was quiet, but the buckets were backwards
Goal
You diagnose and fix exporter output that violates the conventions with the promtool linter, recount the histogram exposition from the raw data, load it into a TSDB with OpenMetrics and check the number of time series, and then build an exporter with only the standard library and actually scrape it.
Why it matters
Prometheus trusts the text an exporter puts out as it is. If a counter lacks _total or units are mixed in milliseconds, dashboards and rules read it with different meanings, and if buckets are not cumulative, quantiles are quietly wrong. You also have no feel for how many time series one label and a few buckets multiply into until you load them. Telling apart what the linter catches and what it cannot is the starting point of an instrumentation review.
Prepared environment
python3 /opt/fixtures/pca_exposition_lab.py init puts legacy.prom (old exporter output that violates the conventions) and payload-sizes.csv (40 request sizes) in /root/pca-exposition/. It actually runs check metrics (lint) and tsdb create-blocks-from openmetrics (block creation) with promtool 3.0.1 from the lab-k8s image. It does not start a Prometheus server. The grader briefly starts the exporters of steps 6 and 7 on a free port and then shuts them down, so the processes you started have nothing to do with grading.
Steps
- After creating the material with
python3 /opt/fixtures/pca_exposition_lab.py init, runpromtool check metrics < /root/pca-exposition/legacy.prom. In/root/pca-exposition/01-lint.txt, writeproblems=(the number of lines flagged),metrics=(the flagged metric names, comma-separated, without duplicates), andmissed_by_lint=(the name of the series that the lint is quiet about but is wrong because its buckets are not cumulative). - Create
/root/pca-exposition/fixed.promand move the five series of legacy.prom into it.http_requestsbecomes the counterhttp_requests_total(keep both samples),request_latency_ms212 becomes the gaugerequest_latency_seconds0.212 (add HELP),queueDepthbecomesqueue_depth,cache_hits_totalbecomes the gaugecache_entries, andjob_duration_countbecomes the gaugebatch_jobs_running. Do not movepayload_bytes.promtool check metricsmust flag nothing. - Using the request sizes in
/root/pca-exposition/payload-sizes.csv, build the histogramhttp_request_size_bytesand append it to the end of fixed.prom. The boundaries are 100, 1000, 10000, and +Inf, and also write_sumand_count. Each bucket is the cumulative count of requests at or below (inclusive of) that boundary. The lint must stay clean. - Move the same samples as fixed.prom into the OpenMetrics format
/root/pca-exposition/scrape.om. Attach a timestamp (in seconds, e.g. 1700000000) to the end of every sample, and the last line is# EOF. After loading withpromtool tsdb create-blocks-from openmetrics /root/pca-exposition/scrape.om /root/pca-exposition/tsdb, writeseries=(NUM SERIES) andsamples=(NUM SAMPLES) in/root/pca-exposition/04-import.txt. - In
/root/pca-exposition/cardinality.om, write thehttp_request_size_byteshistogram with three route label values, /checkout, /search, and /upload, using 12 finite boundaries and +Inf (including _sum and _count for each route, timestamps, and # EOF). Before loading, calculate the number of time series and write it aspredicted_series=in/root/pca-exposition/05-cardinality.txt, and write the load result asimported_series=. - Write
/root/pca-exposition/exporter.pywith only the standard library. It listens on 127.0.0.1 at the environment variablePORT(9464 if not set), and for eachGET /work?d=초(the placeholder is the seconds), it increments the counterdemo_jobs_processed_totalby 1 and records d in the histogramdemo_job_duration_seconds(boundaries 0.1, 0.5, 1, 5).GET /metricsreturns an exposition format with HELP and TYPE, withContent-Type: text/plain; version=0.0.4. The grader starts it fresh on a free port, sends d=0.05, 0.5, 2, and 7, scrapes it, checks the lint, the counter increase, the cumulative buckets, and _sum, and then shuts it down. - Attach a
queuelabel to the counter in exporter.py. Use the name fromGET /work?d=초&queue=이름(the placeholders are the seconds and the name), and if it is absent, usedefault. Escape backslashes, double quotes, and newlines in label values according to the exposition format rules. The grader sends the namesexports,say "hi",c:\tmp, and one containing a newline, and checks that promtool accepts the exposition and that the four names are each read back once with their original values.
Notes
- In the text format, the counter TYPE line uses the name written out through
_total. OpenMetrics distinguishes the series name from the sample name and ends with# EOF. - Common mistakes: writing buckets as per-interval counts, putting a value equal to a boundary into the next bucket, and breaking the line by not escaping a label value.
- The promtool lint rules are the results seen in this version (3.0.1). Other versions may word the findings differently.
- Exposition formats · Metric and label naming · Writing exporters
The six lines promtool rejected
After creating the material with python3 /opt/fixtures/pca_exposition_lab.py init, run promtool check metrics < /root/pca-exposition/legacy.prom. In /root/pca-exposition/01-lint.txt, write problems= (the number of lines flagged), metrics= (the flagged metric names, comma-separated, without duplicates), and missed_by_lint= (the name of the series that the lint is quiet about but is wrong because its buckets are not cumulative).
The linter looks at naming, HELP, and suffix conventions. It does not check whether histogram buckets are cumulative, so read the _bucket values in le order yourself.
Bringing names and units into line with the conventions
Create /root/pca-exposition/fixed.prom and move the five series of legacy.prom into it. http_requests becomes the counter http_requests_total (keep both samples), request_latency_ms 212 becomes the gauge request_latency_seconds 0.212 (add HELP), queueDepth becomes queue_depth, cache_hits_total becomes the gauge cache_entries, and job_duration_count becomes the gauge batch_jobs_running. Do not move payload_bytes. promtool check metrics must flag nothing.
Only counters end in _total. _count, _sum, and _bucket are suffixes used by histograms and summaries. Change units to the base unit (seconds, bytes) and convert the values along with them.
Recounting cumulative buckets from the CSV
Using the request sizes in /root/pca-exposition/payload-sizes.csv, build the histogram http_request_size_bytes and append it to the end of fixed.prom. The boundaries are 100, 1000, 10000, and +Inf, and also write _sum and _count. Each bucket is the cumulative count of requests at or below (inclusive of) that boundary. The lint must stay clean.
The payload_bytes in legacy wrote per-interval counts and was not cumulative. le="+Inf" is always equal to _count.
Loading with OpenMetrics and counting time series
Move the same samples as fixed.prom into the OpenMetrics format /root/pca-exposition/scrape.om. Attach a timestamp (in seconds, e.g. 1700000000) to the end of every sample, and the last line is # EOF. After loading with promtool tsdb create-blocks-from openmetrics /root/pca-exposition/scrape.om /root/pca-exposition/tsdb, write series= (NUM SERIES) and samples= (NUM SAMPLES) in /root/pca-exposition/04-import.txt.
OpenMetrics rejects loading if there is no # EOF or no timestamp. One histogram becomes a time series for each bucket, and _sum and _count each become one too.
The price of one label and 12 buckets
In /root/pca-exposition/cardinality.om, write the http_request_size_bytes histogram with three route label values, /checkout, /search, and /upload, using 12 finite boundaries and +Inf (including _sum and _count for each route, timestamps, and # EOF). Before loading, calculate the number of time series and write it as predicted_series= in /root/pca-exposition/05-cardinality.txt, and write the load result as imported_series=.
The number of time series is the product of label value combinations. One histogram produces (number of boundaries + 1) buckets plus _sum and _count.
Scraping an exporter you built yourself
Write /root/pca-exposition/exporter.py with only the standard library. It listens on 127.0.0.1 at the environment variable PORT (9464 if not set), and for each GET /work?d=초 (the placeholder is the seconds), it increments the counter demo_jobs_processed_total by 1 and records d in the histogram demo_job_duration_seconds (boundaries 0.1, 0.5, 1, 5). GET /metrics returns an exposition format with HELP and TYPE, with Content-Type: text/plain; version=0.0.4. The grader starts it fresh on a free port, sends d=0.05, 0.5, 2, and 7, scrapes it, checks the lint, the counter increase, the cumulative buckets, and _sum, and then shuts it down.
Count buckets cumulatively at or below the boundary. A single request is added to every boundary larger than itself and to +Inf. The grader gives you PORT, so do not fix the port.
A queue name arrived with quotation marks in it
Attach a queue label to the counter in exporter.py. Use the name from GET /work?d=초&queue=이름 (the placeholders are the seconds and the name), and if it is absent, use default. Escape backslashes, double quotes, and newlines in label values according to the exposition format rules. The grader sends the names exports, say "hi", c:\tmp, and one containing a newline, and checks that promtool accepts the exposition and that the four names are each read back once with their original values.
In an exposition format label value, only three characters are escaped. The order matters — if you do not replace the backslash first, the backslashes you newly inserted get replaced again.