TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Why rate Does Not Give You Integers, and Why You Must Not Drop le

Continue in TT Lab

In one line

The two things people get wrong most in PromQL are using rate after sum and dropping le in a quantile calculation. Both return plausible numbers without raising an error, so they survive on dashboards for months.

Why this was needed

Plotted as it is, a counter is a line sloping up to the right with no information. That is why you look at the per-second growth rate with rate. But rate is not simple division; it does extrapolation. It takes the difference between the first and last samples in the range, and then scales it up proportionally for the amount by which the samples are not right against the range boundaries. That is why increase returns values like 3.4. It does so even if there were exactly 3 errors. This is not a bug but the definition, and you do not use it for calculations that need "exactly N."

How it works

Three pitfalls of rate

First, the window must not be narrow compared with the scrape interval. rate produces a value only if there are at least two samples in the window. If the scrape is 15 seconds and the window is 20 seconds, there are moments with only one sample, so the result comes out empty, showing as a hole in the graph and as "condition not met" in an alert. The rule of thumb is window ≥ scrape interval × 4, and for alerts you use 5 minutes or more.

Second, the order.

# 틀림 — 카운터를 먼저 더하면 파드 재시작(리셋)이 감지되지 않는다
rate(sum(http_requests_total) by (route)[5m:])

# 맞음 — rate 를 먼저, 집계는 그다음
sum(rate(http_requests_total[5m])) by (route)

Counter reset correction is correct only when applied to each time series individually. rate and increase treat every moment when a value becomes smaller than the previous sample as a reset, and add the value just before the drop to the increase. If you sum first, this rule is applied to the total rather than to the per-Pod counters, and depending on the case the result goes wrong in two directions.

Third, irate looks at only the last two samples. It is useful on a dashboard for seeing instantaneous response, but you must never use it for alerts. It fires on a single bit of noise.

Lookback delta and staleness

An instant vector selector looks back at most 5 minutes (the default lookback delta) from the evaluation time T and uses the most recent sample. This is why there does not need to be a sample exactly at time T. But if there is a stale marker in between, the time series drops out of the result.

Linear interpolation in histogram_quantile

le="1"     누적 9812
le="2.5"   누적 9993
목표 = 0.99 × 10000 = 9900

추정 = 1.0 + (9900 - 9812) / (9993 - 9812) × (2.5 - 1.0) = 1.729초

A precise-looking number, 1.729 seconds, comes out, but nobody knows where the 181 requests inside this bucket actually are. They could all be at 1.05 seconds or all at 2.4 seconds. Three practical rules come from this. Include a bucket boundary exactly equal to the SLO threshold, place boundaries densely in the range you care about, and set the topmost finite bucket larger than the actual timeout. If p99 exceeds the last finite boundary, the function sticks to that boundary value and you can no longer tell how bad it is.

And when aggregating, always keep le.

# 틀림 — le 를 버리면 버킷이 뭉개져 의미 없는 숫자가 나온다
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (route))

# 맞음
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route))

For the same reason, avg(histogram_quantile(...)) is also wrong. Instead of averaging quantiles, combine the bucket counters first and compute the quantile on top of that.

The rest that the exam likes

Element Key point
offset / @ Relative shift / specifying an absolute epoch time
bool Returns the comparison result as 0/1 values instead of filtering
on / ignoring Specify the matching labels for a binary operation
group_left / group_right Allow many-to-one / one-to-many matching
unless Removes from the left the time series that match the right
Subquery expr[30m:1m] — range:resolution
predict_linear Linear regression extrapolation of a gauge, for capacity alerts
absent Returns 1 when the time series does not exist

What it looks like in the field

The longest-lived bug on my homelab dashboards was a missing by (le. A value comes out, the graph is drawn, and it even moves plausibly. There is only one way to catch it — search for every query that contains histogram_quantile and check by eye whether le is in the aggregation clause. In most teams, at least one is missing.

Cardinality incidents usually show up as a step right after a deployment. If you keep a graph of prometheus_tsdb_head_series up and overlay it with the deployment time, the culprit commit shows up right away. Which metric is the culprit is settled with a single line, topk(10, count by (__name__)({__name__=~".+"})).

What you will do in the next lab

Under /root/pca-promql/, you write eight queries as files. The request rate and error rate per route, a p99 that keeps le, a cardinality diagnosis, a predict_linear disk forecast, an AND gate that prevents phantom alerts at low traffic, a scrape health check that combines up and absent, and finally a burn rate expression that combines two windows with AND. Grading is string matching, so write the specified metric names and windows exactly.