PCA — Prometheus Certified Associate
Writing Eight PromQL Queries
Goal
While writing eight PromQL queries as files, you build the muscle memory for practical disciplines such as the order of rate, preserving le, and a minimum traffic gate. It is the domain with the largest weight on the exam at 28%, and most of the mistakes here are the kind that return a wrong answer without any error.
Why it matters
PromQL is hard not because of the syntax but because it stays quiet even when wrong. If you do not put sum outside rate, the value comes out low; if you leave out by (le), you get an arbitrary number rather than a quantile; and without a minimum traffic gate, phantom alerts fire in every low-traffic stretch before dawn. All three look normal on a dashboard. So queries should be reviewed not by "does it work" but by "under what conditions does it lie." This lab plants those review items in every single query.
Steps
- In
/root/pca-promql/01-request-rate.promql, write the per-second request rate per route. The metric ishttp_requests_total, the window is[5m], and the aggregation label isroute. - In
/root/pca-promql/02-error-ratio.promql, write the 5xx error ratio per route. The numerator is aratefiltered withstatus_class="5xx", the denominator is the totalrate, and you aggregate both withby (route)and divide. - In
/root/pca-promql/03-p99-route.promql, write the p99 latency per route. Insidehistogram_quantile(0.99, ...), apply a[5m]rate tohttp_request_duration_seconds_bucketand aggregate withby (le, route). - In
/root/pca-promql/04-cardinality-topk.promql, write a query that picks the top 10 metrics with the most time series. Usecount by (__name__)insidetopk(10, ...), and the target selector is{__name__=~".+"}. - In
/root/pca-promql/05-disk-predict.promql, write a disk exhaustion forecast. Topredict_linear(node_filesystem_avail_bytes{mountpoint="/data"}[6h], 4 * 3600) < 0, join withandthe condition that the current free ratio is< 0.2(the value ofnode_filesystem_avail_bytesdivided bynode_filesystem_size_bytes). - In
/root/pca-promql/06-nan-guard.promql, write an error-ratio alert expression with a low-traffic guard. The same ratio as in step 2 is> 0.01, and at the same time joinsum(rate(http_requests_total[5m])) by (route) > 1withand. - In
/root/pca-promql/07-scrape-health.promql, write a scrape health check. Joinup{job="checkout-api"} == 0andabsent(up{job="checkout-api"})withor. - In
/root/pca-promql/08-burn-rate.promql, write a multi-window burn rate. Join the three termsjob:slo_errors:ratio_rate1h{job="checkout-api"} > (14.4 * 0.001),job:slo_errors:ratio_rate5m{job="checkout-api"} > (14.4 * 0.001), andjob:http_requests:rate5m{job="checkout-api"} > 1withand.
Notes
- Grading ignores whitespace and
#comments and compares strings. You may write it prettily over several lines, but write the metric names, labels, windows, and numbers exactly as instructed. - Common mistake 1: the
rate(sum(...))order. Grading explicitly rejects this form. - Common mistake 2: using only
by (route)in step 3 and leaving outle. - Common mistake 3: attaching
by (route)only to the numerator in step 2. For division, the label sets on both sides must match.
Per-second request rate per route
In /root/pca-promql/01-request-rate.promql, write the per-second request rate per route. The metric is http_requests_total, the window is [5m], and the aggregation label is route.
Plotted as it is, a counter is a line sloping up to the right. Apply rate first and then aggregate that result. If you do it the other way around, Pod restart resets are not corrected.
5xx error ratio per route
In /root/pca-promql/02-error-ratio.promql, write the 5xx error ratio per route. The numerator is a rate filtered with status_class="5xx", the denominator is the total rate, and you aggregate both with by (route) and divide.
Since it is a ratio, you need a numerator and a denominator separately. To divide two vectors, their label sets must match, so both sides must be aggregated along the same dimension.
p99 latency per route
In /root/pca-promql/03-p99-route.promql, write the p99 latency per route. Inside histogram_quantile(0.99, ...), apply a [5m] rate to http_request_duration_seconds_bucket and aggregate with by (le, route).
Buckets are counters, so apply rate first. If you lose le in the aggregation, the buckets get smeared and a meaningless number comes out without any error. Never use a form that averages quantiles.
Diagnosing the top 10 by cardinality
In /root/pca-promql/04-cardinality-topk.promql, write a query that picks the top 10 metrics with the most time series. Use count by (__name__) inside topk(10, ...), and the target selector is {__name__=~".+"}.
To count how many time series each metric name has, group count by name. The selector that picks every metric applies a regular expression to the name label. Keep only the top with topk.
Disk exhaustion forecast
In /root/pca-promql/05-disk-predict.promql, write a disk exhaustion forecast. To predict_linear(node_filesystem_avail_bytes{mountpoint="/data"}[6h], 4 * 3600) < 0, join with and the condition that the current free ratio is < 0.2 (the value of node_filesystem_avail_bytes divided by node_filesystem_size_bytes).
predict_linear extrapolates a gauge's recent trend by linear regression. The second argument is in seconds. Looking only at the trend can fire even on a disk with plenty of room, so also attach the current free ratio condition with AND.
Preventing low-traffic phantom alerts
In /root/pca-promql/06-nan-guard.promql, write an error-ratio alert expression with a low-traffic guard. The same ratio as in step 2 is > 0.01, and at the same time join sum(rate(http_requests_total[5m])) by (route) > 1 with and.
In a period when 3 requests come in over 5 minutes, if 1 fails the error ratio is 33%. With the ratio condition alone, an alert fires every night before dawn. Attach a minimum traffic condition with AND. If there are 0 requests at all, the ratio becomes NaN and the condition quietly disappears.
Scrape health check
In /root/pca-promql/07-scrape-health.promql, write a scrape health check. Join up{job="checkout-api"} == 0 and absent(up{job="checkout-api"}) with or.
A situation where up is 0 and a situation where the up time series itself has vanished are different. The latter is never caught by a value comparison. Join with OR the function that returns 1 when the time series does not exist.
Multi-window burn rate
In /root/pca-promql/08-burn-rate.promql, write a multi-window burn rate. Join the three terms job:slo_errors:ratio_rate1h{job="checkout-api"} > (14.4 * 0.001), job:slo_errors:ratio_rate5m{job="checkout-api"} > (14.4 * 0.001), and job:http_requests:rate5m{job="checkout-api"} > 1 with and.
The long window decides "are we burning at this rate," and the short window decides "is it still going on now." Without the short window, the alert lingers for the length of the long window even after the incident is over. Add to this a minimum traffic gate and join the three terms with AND. Reference the recording rule time series names exactly.