TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

The batch stopped, and the alert resolved itself

Continue in TT Lab

Goal

You evaluate, on a fixed 240-minute time series, the difference between a metric that holds a time as its value and timestamp(), staleness markers and lookback, and *_over_time, subqueries, and deriv, and confirm them with numbers, and you build an alerting rule that does not clear even when the target disappears.

Why it matters

A freshness alert for a batch job is often silent in the biggest incident. When a Pod disappears, its time series disappears too, and a condition expression on a vanished time series becomes an empty result, neither true nor false. You can trust why an alert is quiet only if you know when PromQL sees samples and when it does not (lookback, staleness), and what instantaneous values and time aggregations hide.

Prepared environment

python3 /opt/fixtures/pca_staleness_lab.py init puts the story of the time series in /root/pca-staleness/README.txt. You can see the input notation with python3 /opt/fixtures/pca_staleness_lab.py series. The times run from minute 0 to minute 240 at 1-minute intervals, and because it is a promtool test time, time() is the evaluation time in seconds (9000 at 150m).

python3 /opt/fixtures/pca_staleness_lab.py eval -f 파일 --at 150m (the placeholder is the file) evaluates with the promtool 3.0.1 engine from lab-k8s. The grader recomputes your expressions and the reference expressions on the same time series, and in the last step it runs your rule file through a rule test the grader builds.

Steps

  1. In /root/pca-staleness/01-age.promql, write the seconds elapsed since nightly-export's last success using time() and batch_last_success_timestamp_seconds. Write the result of python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m as age_seconds= in /root/pca-staleness/01-age.txt. The material is created by python3 /opt/fixtures/pca_staleness_lab.py init.
  2. In /root/pca-staleness/02-wrong.promql, write time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}) and evaluate it at 150m. In /root/pca-staleness/02-trap.txt, write wrong_age_seconds= (the result of this expression), sample_timestamp= (the result of timestamp()), and value= (the metric value).
  3. Evaluate the batch_last_success_timestamp_seconds of the two jobs minute by minute starting at 178m, and find the last minute at which a result comes out. In /root/pca-staleness/03-lookback.txt, write as integers nightly_last_minute= (the job into which a staleness marker was inserted) and federated_last_minute= (the job whose samples simply stopped, without a marker).
  4. Evaluate time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600 at 170m and at 200m, and write the number of result time series as naive_series_170m= and naive_series_200m= in /root/pca-staleness/04-alert.txt. Then, in /root/pca-staleness/04-fixed.promql, write an expression that still returns a result after the target has disappeared. The grader checks that there is an empty result at 60m and 125m and a result at 135m, 170m, 182m, 200m, and 235m.
  5. In /root/pca-staleness/05-availability.promql, write the availability of job="api" over the last 30 minutes with avg_over_time. Evaluate it at 120m and write availability_30m=, instant_up= (the up{job="api"} at the same time), and min_up_30m= (the result of min_over_time) in /root/pca-staleness/05-availability.txt.
  6. In /root/pca-staleness/06-slope.promql, write the 10-minute slope (per second) of queue_depth{queue="exports"} with deriv. In /root/pca-staleness/06-slope.txt, write deriv_110m= (the result at 110m), rate_125m= (rate(queue_depth[10m]) at 125m), and deriv_125m= (deriv at 125m). By 120m the queue had mostly been processed.
  7. In /root/pca-staleness/07-peak.promql, use a subquery to write the maximum of rate(batch_rows_processed_total[5m]) viewed at 1-minute intervals over the last hour. Evaluate it at 80m and write peak_rows_per_sec= and hour_avg_rows_per_sec= (the result of rate(...[1h])) in /root/pca-staleness/07-peak.txt.
  8. In /root/pca-staleness/08-rules.yml, write one group and the alert BatchExportStale. expr is an expression that produces a result even when the target disappears, as in step 04, for: 10m, and labels.severity: ticket. Check the syntax with promtool check rules. The grader builds a rule test on the same time series and checks that no alert has fired at 60m, 125m, and 135m (pending), and that an alert with job="nightly-export", severity="ticket" is firing at 145m, 150m, 200m, and 235m.

Notes

How many seconds ago was the last success

In /root/pca-staleness/01-age.promql, write the seconds elapsed since nightly-export's last success using time() and batch_last_success_timestamp_seconds. Write the result of python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m as age_seconds= in /root/pca-staleness/01-age.txt. The material is created by python3 /opt/fixtures/pca_staleness_lab.py init.

The metric's value itself is a Unix time (in seconds). Subtracting that value from the evaluation time gives the elapsed time.

Measure with timestamp() and it is always fresh

In /root/pca-staleness/02-wrong.promql, write time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}) and evaluate it at 150m. In /root/pca-staleness/02-trap.txt, write wrong_age_seconds= (the result of this expression), sample_timestamp= (the result of timestamp()), and value= (the metric value).

timestamp() is the time the sample was scraped. While it is scraped every minute, that time keeps getting newer.

Until when is a vanished time series visible

Evaluate the batch_last_success_timestamp_seconds of the two jobs minute by minute starting at 178m, and find the last minute at which a result comes out. In /root/pca-staleness/03-lookback.txt, write as integers nightly_last_minute= (the job into which a staleness marker was inserted) and federated_last_minute= (the job whose samples simply stopped, without a marker).

When a target disappears, Prometheus inserts a staleness marker and removes it immediately. Without a marker, it keeps returning the last sample for the lookback (5 minutes by default). Also check whether the range boundary is included.

The alert cleared once the target disappeared

Evaluate time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600 at 170m and at 200m, and write the number of result time series as naive_series_170m= and naive_series_200m= in /root/pca-staleness/04-alert.txt. Then, in /root/pca-staleness/04-fixed.promql, write an expression that still returns a result after the target has disappeared. The grader checks that there is an empty result at 60m and 125m and a result at 135m, 170m, 182m, 200m, and 235m.

A vanished time series has no value to compare, so the condition expression becomes an empty result. Attach with or the function that turns absence into a signal.

Up=1 right now, and the 30-minute availability is

In /root/pca-staleness/05-availability.promql, write the availability of job="api" over the last 30 minutes with avg_over_time. Evaluate it at 120m and write availability_30m=, instant_up= (the up{job="api"} at the same time), and min_up_30m= (the result of min_over_time) in /root/pca-staleness/05-availability.txt.

The time average of a gauge of 0s and 1s is the fraction of samples that were 1. An instantaneous value does not show an outage inside the window.

When you apply rate to a gauge

In /root/pca-staleness/06-slope.promql, write the 10-minute slope (per second) of queue_depth{queue="exports"} with deriv. In /root/pca-staleness/06-slope.txt, write deriv_110m= (the result at 110m), rate_125m= (rate(queue_depth[10m]) at 125m), and deriv_125m= (deriv at 125m). By 120m the queue had mostly been processed.

rate treats a drop in value as a counter restart and corrects for it. A decrease in a gauge is a real change, not a restart.

The peak throughput hidden in a one-hour average

In /root/pca-staleness/07-peak.promql, use a subquery to write the maximum of rate(batch_rows_processed_total[5m]) viewed at 1-minute intervals over the last hour. Evaluate it at 80m and write peak_rows_per_sec= and hour_avg_rows_per_sec= (the result of rate(...[1h])) in /root/pca-staleness/07-peak.txt.

식[범위:간격] (the placeholders are the expression, the range, and the interval) re-evaluates the inner expression at each interval to make a range vector. You can apply *_over_time on top of that.

An alerting rule that keeps firing even when the target vanishes

In /root/pca-staleness/08-rules.yml, write one group and the alert BatchExportStale. expr is an expression that produces a result even when the target disappears, as in step 04, for: 10m, and labels.severity: ticket. Check the syntax with promtool check rules. The grader builds a rule test on the same time series and checks that no alert has fired at 60m, 125m, and 135m (pending), and that an alert with job="nightly-export", severity="ticket" is firing at 145m, 150m, 200m, and 235m.

absent() attaches the labels of its equality matchers to the result. If the labels of the two branches are the same, the alert continues without breaking. for is the wait time before firing.