TT Lab
开始
学习 学习路径 课程

PCA — Prometheus 认证助理

批处理停了,告警却自己恢复了

在 TT Lab 中继续学习

目标

在固定的 240 分钟时间序列上,对以时间戳为值的指标与 timestamp() 的差异、staleness marker 与 lookback、*_over_time、子查询和 deriv 进行求值并用数字加以确认,然后做出一条在目标消失后也不会自行解除的告警规则。

为什么重要

批处理作业的新鲜度告警,往往会在最大的事故面前保持沉默。Pod 消失后,时间序列也随之消失,而对已消失的时间序列设置的条件表达式,既不是真也不是假,而是空结果。只有了解 PromQL 何时读取样本、何时不读取(lookback、staleness),以及瞬时值和按时间聚合各自隐藏了什么,才能相信告警沉默的原因。

已准备的环境

python3 /opt/fixtures/pca_staleness_lab.py init 会在 /root/pca-staleness/README.txt 中放入这组时间序列的背景说明。用 python3 /opt/fixtures/pca_staleness_lab.py series 可以查看输入的表示法。时刻从第 0 分钟到第 240 分钟,每 1 分钟一个点;这是 promtool 测试用的时间,所以 time() 是求值时刻的秒数(150m 时为 9000)。

python3 /opt/fixtures/pca_staleness_lab.py eval -f 파일 --at 150m(占位符为文件名)使用 lab-k8s 的 promtool 3.0.1 引擎求值。评分器用同样的时间序列重新计算你的表达式和基准表达式,最后一步则把你的规则文件放进评分器生成的规则测试中运行。

步骤

  1. 在 /root/pca-staleness/01-age.promql 中,用 time() 和 batch_last_success_timestamp_seconds 写出 nightly-export 自上次成功以来经过的秒数。把 python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m 的结果以 age_seconds= 写入 /root/pca-staleness/01-age.txt。材料由 python3 /opt/fixtures/pca_staleness_lab.py init 生成。
  2. 在 /root/pca-staleness/02-wrong.promql 中写入 time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}),并在 150m 处求值。在 /root/pca-staleness/02-trap.txt 中写下 wrong_age_seconds=(这个表达式的结果)、sample_timestamp=(timestamp() 的结果)和 value=(指标值)。
  3. 从 178m 起每隔 1 分钟对两个 job 的 batch_last_success_timestamp_seconds 求值,找出仍有结果的最后一分钟。在 /root/pca-staleness/03-lookback.txt 中用整数写下 nightly_last_minute=(写入了 staleness marker 的 job)和 federated_last_minute=(没有 marker、只是样本中断的 job)。
  4. 在 170m 和 200m 处对 time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600 求值,把结果时间序列的数量写到 /root/pca-staleness/04-alert.txt 的 naive_series_170m= 和 naive_series_200m= 中。然后在 /root/pca-staleness/04-fixed.promql 中写出一个在目标消失之后仍会返回结果的表达式。评分器会检查 60m、125m 处是否为空结果,135m、170m、182m、200m、235m 处是否有结果。
  5. 在 /root/pca-staleness/05-availability.promql 中,用 avg_over_time 写出 job="api" 最近 30 分钟的可用率。在 120m 处求值,并在 /root/pca-staleness/05-availability.txt 中写下 availability_30m=、instant_up=(同一时刻的 up{job="api"})和 min_up_30m=(min_over_time 的结果)。
  6. 在 /root/pca-staleness/06-slope.promql 中,用 deriv 写出 queue_depth{queue="exports"} 的 10 分钟斜率(每秒)。在 /root/pca-staleness/06-slope.txt 中写下 deriv_110m=(110m 的结果)、rate_125m=(125m 处的 rate(queue_depth[10m]))和 deriv_125m=(125m 处的 deriv)。120m 时队列已大部分处理完毕。
  7. 在 /root/pca-staleness/07-peak.promql 中,用子查询写出最近 1 小时内、每隔 1 分钟观察到的 rate(batch_rows_processed_total[5m]) 的最大值。在 80m 处求值,并在 /root/pca-staleness/07-peak.txt 中写下 peak_rows_per_sec= 和 hour_avg_rows_per_sec=(rate(...[1h]) 的结果)。
  8. 在 /root/pca-staleness/08-rules.yml 中写一个分组和告警 BatchExportStale。expr 使用像第 4 步那样在目标消失后仍有结果的表达式,for: 10m,labels.severity: ticket。用 promtool check rules 检查语法。评分器会用同样的时间序列生成规则测试,确认 60m、125m、135m(pending)时没有触发的告警,而 145m、150m、200m、235m 时 job="nightly-export"、severity="ticket" 的告警处于触发状态。

参考

上次成功是多少秒之前

在 /root/pca-staleness/01-age.promql 中,用 time() 和 batch_last_success_timestamp_seconds 写出 nightly-export 自上次成功以来经过的秒数。把 python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m 的结果以 age_seconds= 写入 /root/pca-staleness/01-age.txt。材料由 python3 /opt/fixtures/pca_staleness_lab.py init 生成。

指标的值本身就是 Unix 时间(秒)。用求值时刻减去这个值,就是经过的时间。

用 timestamp() 来测,永远是新鲜的

在 /root/pca-staleness/02-wrong.promql 中写入 time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}),并在 150m 处求值。在 /root/pca-staleness/02-trap.txt 中写下 wrong_age_seconds=(这个表达式的结果)、sample_timestamp=(timestamp() 的结果)和 value=(指标值)。

timestamp() 是抓取样本的时刻。只要每分钟都在抓取,这个时刻就会一直更新。

消失的时间序列会被看到多久

从 178m 起每隔 1 分钟对两个 job 的 batch_last_success_timestamp_seconds 求值,找出仍有结果的最后一分钟。在 /root/pca-staleness/03-lookback.txt 中用整数写下 nightly_last_minute=(写入了 staleness marker 的 job)和 federated_last_minute=(没有 marker、只是样本中断的 job)。

目标消失时,Prometheus 会写入 staleness marker 并让它立即消失。没有 marker 时,会在 lookback(默认 5 分钟)内持续返回最后一个样本。也要确认区间边界是否包含在内。

目标消失后,告警解除了

在 170m 和 200m 处对 time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600 求值,把结果时间序列的数量写到 /root/pca-staleness/04-alert.txt 的 naive_series_170m= 和 naive_series_200m= 中。然后在 /root/pca-staleness/04-fixed.promql 中写出一个在目标消失之后仍会返回结果的表达式。评分器会检查 60m、125m 处是否为空结果,135m、170m、182m、200m、235m 处是否有结果。

已消失的时间序列没有可比较的值,所以条件表达式成为空结果。用 or 接上一个能把“不存在”变成信号的函数。

现在 up=1,30 分钟可用率是多少

在 /root/pca-staleness/05-availability.promql 中,用 avg_over_time 写出 job="api" 最近 30 分钟的可用率。在 120m 处求值,并在 /root/pca-staleness/05-availability.txt 中写下 availability_30m=、instant_up=(同一时刻的 up{job="api"})和 min_up_30m=(min_over_time 的结果)。

由 0 和 1 组成的仪表的时间平均,就是值为 1 的样本所占的比例。瞬时值看不出窗口内的故障。

对仪表使用 rate 会怎样

在 /root/pca-staleness/06-slope.promql 中,用 deriv 写出 queue_depth{queue="exports"} 的 10 分钟斜率(每秒)。在 /root/pca-staleness/06-slope.txt 中写下 deriv_110m=(110m 的结果)、rate_125m=(125m 处的 rate(queue_depth[10m]))和 deriv_125m=(125m 处的 deriv)。120m 时队列已大部分处理完毕。

值下降时,rate 会认为计数器重启了并加以修正。仪表的下降不是重启,而是真实的变化。

隐藏在一小时平均值里的峰值吞吐量

在 /root/pca-staleness/07-peak.promql 中,用子查询写出最近 1 小时内、每隔 1 分钟观察到的 rate(batch_rows_processed_total[5m]) 的最大值。在 80m 处求值,并在 /root/pca-staleness/07-peak.txt 中写下 peak_rows_per_sec= 和 hour_avg_rows_per_sec=(rate(...[1h]) 的结果)。

식[범위:간격](占位符依次为表达式、范围和间隔)会每隔一个间隔重新对内层表达式求值,生成范围向量。之后就可以在其上使用 *_over_time。

目标消失后仍会持续响起的告警规则

在 /root/pca-staleness/08-rules.yml 中写一个分组和告警 BatchExportStale。expr 使用像第 4 步那样在目标消失后仍有结果的表达式,for: 10m,labels.severity: ticket。用 promtool check rules 检查语法。评分器会用同样的时间序列生成规则测试,确认 60m、125m、135m(pending)时没有触发的告警,而 145m、150m、200m、235m 时 job="nightly-export"、severity="ticket" 的告警处于触发状态。

absent() 会把等号匹配器中的标签附加到结果上。两个分支的标签相同,告警就不会中断而是持续下去。for 是触发前的等待时间。