PCA — Prometheus Certified Associate
The batch stopped, and the alert resolved itself
Goal
You evaluate, on a fixed 240-minute time series, the difference between a metric that holds a time as its value and timestamp(), staleness markers and lookback, and *_over_time, subqueries, and deriv, and confirm them with numbers, and you build an alerting rule that does not clear even when the target disappears.
Why it matters
A freshness alert for a batch job is often silent in the biggest incident. When a Pod disappears, its time series disappears too, and a condition expression on a vanished time series becomes an empty result, neither true nor false. You can trust why an alert is quiet only if you know when PromQL sees samples and when it does not (lookback, staleness), and what instantaneous values and time aggregations hide.
Prepared environment
python3 /opt/fixtures/pca_staleness_lab.py init puts the story of the time series in /root/pca-staleness/README.txt. You can see the input notation with python3 /opt/fixtures/pca_staleness_lab.py series. The times run from minute 0 to minute 240 at 1-minute intervals, and because it is a promtool test time, time() is the evaluation time in seconds (9000 at 150m).
python3 /opt/fixtures/pca_staleness_lab.py eval -f 파일 --at 150m (the placeholder is the file) evaluates with the promtool 3.0.1 engine from lab-k8s. The grader recomputes your expressions and the reference expressions on the same time series, and in the last step it runs your rule file through a rule test the grader builds.
Steps
- In
/root/pca-staleness/01-age.promql, write the seconds elapsed since nightly-export's last success usingtime()andbatch_last_success_timestamp_seconds. Write the result ofpython3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150masage_seconds=in/root/pca-staleness/01-age.txt. The material is created bypython3 /opt/fixtures/pca_staleness_lab.py init. - In
/root/pca-staleness/02-wrong.promql, writetime() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"})and evaluate it at 150m. In/root/pca-staleness/02-trap.txt, writewrong_age_seconds=(the result of this expression),sample_timestamp=(the result oftimestamp()), andvalue=(the metric value). - Evaluate the
batch_last_success_timestamp_secondsof the two jobs minute by minute starting at 178m, and find the last minute at which a result comes out. In/root/pca-staleness/03-lookback.txt, write as integersnightly_last_minute=(the job into which a staleness marker was inserted) andfederated_last_minute=(the job whose samples simply stopped, without a marker). - Evaluate
time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600at 170m and at 200m, and write the number of result time series asnaive_series_170m=andnaive_series_200m=in/root/pca-staleness/04-alert.txt. Then, in/root/pca-staleness/04-fixed.promql, write an expression that still returns a result after the target has disappeared. The grader checks that there is an empty result at 60m and 125m and a result at 135m, 170m, 182m, 200m, and 235m. - In
/root/pca-staleness/05-availability.promql, write the availability of job="api" over the last 30 minutes withavg_over_time. Evaluate it at 120m and writeavailability_30m=,instant_up=(theup{job="api"}at the same time), andmin_up_30m=(the result ofmin_over_time) in/root/pca-staleness/05-availability.txt. - In
/root/pca-staleness/06-slope.promql, write the 10-minute slope (per second) ofqueue_depth{queue="exports"}withderiv. In/root/pca-staleness/06-slope.txt, writederiv_110m=(the result at 110m),rate_125m=(rate(queue_depth[10m])at 125m), andderiv_125m=(derivat 125m). By 120m the queue had mostly been processed. - In
/root/pca-staleness/07-peak.promql, use a subquery to write the maximum ofrate(batch_rows_processed_total[5m])viewed at 1-minute intervals over the last hour. Evaluate it at 80m and writepeak_rows_per_sec=andhour_avg_rows_per_sec=(the result ofrate(...[1h])) in/root/pca-staleness/07-peak.txt. - In
/root/pca-staleness/08-rules.yml, write one group and the alertBatchExportStale.expris an expression that produces a result even when the target disappears, as in step 04,for: 10m, andlabels.severity: ticket. Check the syntax withpromtool check rules. The grader builds a rule test on the same time series and checks that no alert has fired at 60m, 125m, and 135m (pending), and that an alert with job="nightly-export", severity="ticket" is firing at 145m, 150m, 200m, and 235m.
Notes
- Range selection is open on the left.
[5m]does not include the sample exactly 5 minutes before the evaluation time. absent(v)returns a single time series with the value 1 when v does not exist, and an empty result when it does.- Common mistakes: measuring elapsed time with
timestamp(), applyingrateto a gauge, and trying to catch a vanished target withup == 0alone. - This time series is synthetic material for promtool tests. It does not reproduce a real server's scrape delays or retries.
- Querying basics — Staleness · Functions · Unit testing for rules
How many seconds ago was the last success
In /root/pca-staleness/01-age.promql, write the seconds elapsed since nightly-export's last success using time() and batch_last_success_timestamp_seconds. Write the result of python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m as age_seconds= in /root/pca-staleness/01-age.txt. The material is created by python3 /opt/fixtures/pca_staleness_lab.py init.
The metric's value itself is a Unix time (in seconds). Subtracting that value from the evaluation time gives the elapsed time.
Measure with timestamp() and it is always fresh
In /root/pca-staleness/02-wrong.promql, write time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}) and evaluate it at 150m. In /root/pca-staleness/02-trap.txt, write wrong_age_seconds= (the result of this expression), sample_timestamp= (the result of timestamp()), and value= (the metric value).
timestamp() is the time the sample was scraped. While it is scraped every minute, that time keeps getting newer.
Until when is a vanished time series visible
Evaluate the batch_last_success_timestamp_seconds of the two jobs minute by minute starting at 178m, and find the last minute at which a result comes out. In /root/pca-staleness/03-lookback.txt, write as integers nightly_last_minute= (the job into which a staleness marker was inserted) and federated_last_minute= (the job whose samples simply stopped, without a marker).
When a target disappears, Prometheus inserts a staleness marker and removes it immediately. Without a marker, it keeps returning the last sample for the lookback (5 minutes by default). Also check whether the range boundary is included.
The alert cleared once the target disappeared
Evaluate time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600 at 170m and at 200m, and write the number of result time series as naive_series_170m= and naive_series_200m= in /root/pca-staleness/04-alert.txt. Then, in /root/pca-staleness/04-fixed.promql, write an expression that still returns a result after the target has disappeared. The grader checks that there is an empty result at 60m and 125m and a result at 135m, 170m, 182m, 200m, and 235m.
A vanished time series has no value to compare, so the condition expression becomes an empty result. Attach with or the function that turns absence into a signal.
Up=1 right now, and the 30-minute availability is
In /root/pca-staleness/05-availability.promql, write the availability of job="api" over the last 30 minutes with avg_over_time. Evaluate it at 120m and write availability_30m=, instant_up= (the up{job="api"} at the same time), and min_up_30m= (the result of min_over_time) in /root/pca-staleness/05-availability.txt.
The time average of a gauge of 0s and 1s is the fraction of samples that were 1. An instantaneous value does not show an outage inside the window.
When you apply rate to a gauge
In /root/pca-staleness/06-slope.promql, write the 10-minute slope (per second) of queue_depth{queue="exports"} with deriv. In /root/pca-staleness/06-slope.txt, write deriv_110m= (the result at 110m), rate_125m= (rate(queue_depth[10m]) at 125m), and deriv_125m= (deriv at 125m). By 120m the queue had mostly been processed.
rate treats a drop in value as a counter restart and corrects for it. A decrease in a gauge is a real change, not a restart.
The peak throughput hidden in a one-hour average
In /root/pca-staleness/07-peak.promql, use a subquery to write the maximum of rate(batch_rows_processed_total[5m]) viewed at 1-minute intervals over the last hour. Evaluate it at 80m and write peak_rows_per_sec= and hour_avg_rows_per_sec= (the result of rate(...[1h])) in /root/pca-staleness/07-peak.txt.
식[범위:간격] (the placeholders are the expression, the range, and the interval) re-evaluates the inner expression at each interval to make a range vector. You can apply *_over_time on top of that.
An alerting rule that keeps firing even when the target vanishes
In /root/pca-staleness/08-rules.yml, write one group and the alert BatchExportStale. expr is an expression that produces a result even when the target disappears, as in step 04, for: 10m, and labels.severity: ticket. Check the syntax with promtool check rules. The grader builds a rule test on the same time series and checks that no alert has fired at 60m, 125m, and 135m (pending), and that an alert with job="nightly-export", severity="ticket" is firing at 145m, 150m, 200m, and 235m.
absent() attaches the labels of its equality matchers to the result. If the labels of the two branches are the same, the alert continues without breaking. for is the wait time before firing.