批处理停了,告警却自己恢复了
目标
在固定的 240 分钟时间序列上,对以时间戳为值的指标与 timestamp() 的差异、staleness marker 与 lookback、*_over_time、子查询和 deriv 进行求值并用数字加以确认,然后做出一条在目标消失后也不会自行解除的告警规则。
为什么重要
批处理作业的新鲜度告警,往往会在最大的事故面前保持沉默。Pod 消失后,时间序列也随之消失,而对已消失的时间序列设置的条件表达式,既不是真也不是假,而是空结果。只有了解 PromQL 何时读取样本、何时不读取(lookback、staleness),以及瞬时值和按时间聚合各自隐藏了什么,才能相信告警沉默的原因。
已准备的环境
python3 /opt/fixtures/pca_staleness_lab.py init 会在 /root/pca-staleness/README.txt 中放入这组时间序列的背景说明。用 python3 /opt/fixtures/pca_staleness_lab.py series 可以查看输入的表示法。时刻从第 0 分钟到第 240 分钟,每 1 分钟一个点;这是 promtool 测试用的时间,所以 time() 是求值时刻的秒数(150m 时为 9000)。
python3 /opt/fixtures/pca_staleness_lab.py eval -f 파일 --at 150m(占位符为文件名)使用 lab-k8s 的 promtool 3.0.1 引擎求值。评分器用同样的时间序列重新计算你的表达式和基准表达式,最后一步则把你的规则文件放进评分器生成的规则测试中运行。
步骤
- 在
/root/pca-staleness/01-age.promql中,用time()和batch_last_success_timestamp_seconds写出 nightly-export 自上次成功以来经过的秒数。把python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m的结果以age_seconds=写入/root/pca-staleness/01-age.txt。材料由python3 /opt/fixtures/pca_staleness_lab.py init生成。 - 在
/root/pca-staleness/02-wrong.promql中写入time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}),并在 150m 处求值。在/root/pca-staleness/02-trap.txt中写下wrong_age_seconds=(这个表达式的结果)、sample_timestamp=(timestamp()的结果)和value=(指标值)。 - 从 178m 起每隔 1 分钟对两个 job 的
batch_last_success_timestamp_seconds求值,找出仍有结果的最后一分钟。在/root/pca-staleness/03-lookback.txt中用整数写下nightly_last_minute=(写入了 staleness marker 的 job)和federated_last_minute=(没有 marker、只是样本中断的 job)。 - 在 170m 和 200m 处对
time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600求值,把结果时间序列的数量写到/root/pca-staleness/04-alert.txt的naive_series_170m=和naive_series_200m=中。然后在/root/pca-staleness/04-fixed.promql中写出一个在目标消失之后仍会返回结果的表达式。评分器会检查 60m、125m 处是否为空结果,135m、170m、182m、200m、235m 处是否有结果。 - 在
/root/pca-staleness/05-availability.promql中,用avg_over_time写出 job="api" 最近 30 分钟的可用率。在 120m 处求值,并在/root/pca-staleness/05-availability.txt中写下availability_30m=、instant_up=(同一时刻的up{job="api"})和min_up_30m=(min_over_time的结果)。 - 在
/root/pca-staleness/06-slope.promql中,用deriv写出queue_depth{queue="exports"}的 10 分钟斜率(每秒)。在/root/pca-staleness/06-slope.txt中写下deriv_110m=(110m 的结果)、rate_125m=(125m 处的rate(queue_depth[10m]))和deriv_125m=(125m 处的deriv)。120m 时队列已大部分处理完毕。 - 在
/root/pca-staleness/07-peak.promql中,用子查询写出最近 1 小时内、每隔 1 分钟观察到的rate(batch_rows_processed_total[5m])的最大值。在 80m 处求值,并在/root/pca-staleness/07-peak.txt中写下peak_rows_per_sec=和hour_avg_rows_per_sec=(rate(...[1h])的结果)。 - 在
/root/pca-staleness/08-rules.yml中写一个分组和告警BatchExportStale。expr使用像第 4 步那样在目标消失后仍有结果的表达式,for: 10m,labels.severity: ticket。用promtool check rules检查语法。评分器会用同样的时间序列生成规则测试,确认 60m、125m、135m(pending)时没有触发的告警,而 145m、150m、200m、235m 时 job="nightly-export"、severity="ticket" 的告警处于触发状态。
参考
- 范围选择器的左侧是开区间。
[5m]不包含求值时刻 5 分钟前的那个样本。 absent(v)在 v 不存在时返回一条值为 1 的时间序列,存在时返回空结果。- 常见错误:用
timestamp()来衡量经过的时间;对仪表使用rate;只靠up == 0就想抓住已经消失的目标。 - 这组时间序列是用于 promtool 测试的合成数据,不会重现真实服务器的抓取延迟和重试。
- Querying basics — Staleness、Functions、Unit testing for rules
上次成功是多少秒之前
在 /root/pca-staleness/01-age.promql 中,用 time() 和 batch_last_success_timestamp_seconds 写出 nightly-export 自上次成功以来经过的秒数。把 python3 /opt/fixtures/pca_staleness_lab.py eval -f /root/pca-staleness/01-age.promql --at 150m 的结果以 age_seconds= 写入 /root/pca-staleness/01-age.txt。材料由 python3 /opt/fixtures/pca_staleness_lab.py init 生成。
指标的值本身就是 Unix 时间(秒)。用求值时刻减去这个值,就是经过的时间。
用 timestamp() 来测,永远是新鲜的
在 /root/pca-staleness/02-wrong.promql 中写入 time() - timestamp(batch_last_success_timestamp_seconds{job="nightly-export"}),并在 150m 处求值。在 /root/pca-staleness/02-trap.txt 中写下 wrong_age_seconds=(这个表达式的结果)、sample_timestamp=(timestamp() 的结果)和 value=(指标值)。
timestamp() 是抓取样本的时刻。只要每分钟都在抓取,这个时刻就会一直更新。
消失的时间序列会被看到多久
从 178m 起每隔 1 分钟对两个 job 的 batch_last_success_timestamp_seconds 求值,找出仍有结果的最后一分钟。在 /root/pca-staleness/03-lookback.txt 中用整数写下 nightly_last_minute=(写入了 staleness marker 的 job)和 federated_last_minute=(没有 marker、只是样本中断的 job)。
目标消失时,Prometheus 会写入 staleness marker 并让它立即消失。没有 marker 时,会在 lookback(默认 5 分钟)内持续返回最后一个样本。也要确认区间边界是否包含在内。
目标消失后,告警解除了
在 170m 和 200m 处对 time() - batch_last_success_timestamp_seconds{job="nightly-export"} > 3600 求值,把结果时间序列的数量写到 /root/pca-staleness/04-alert.txt 的 naive_series_170m= 和 naive_series_200m= 中。然后在 /root/pca-staleness/04-fixed.promql 中写出一个在目标消失之后仍会返回结果的表达式。评分器会检查 60m、125m 处是否为空结果,135m、170m、182m、200m、235m 处是否有结果。
已消失的时间序列没有可比较的值,所以条件表达式成为空结果。用 or 接上一个能把“不存在”变成信号的函数。
现在 up=1,30 分钟可用率是多少
在 /root/pca-staleness/05-availability.promql 中,用 avg_over_time 写出 job="api" 最近 30 分钟的可用率。在 120m 处求值,并在 /root/pca-staleness/05-availability.txt 中写下 availability_30m=、instant_up=(同一时刻的 up{job="api"})和 min_up_30m=(min_over_time 的结果)。
由 0 和 1 组成的仪表的时间平均,就是值为 1 的样本所占的比例。瞬时值看不出窗口内的故障。
对仪表使用 rate 会怎样
在 /root/pca-staleness/06-slope.promql 中,用 deriv 写出 queue_depth{queue="exports"} 的 10 分钟斜率(每秒)。在 /root/pca-staleness/06-slope.txt 中写下 deriv_110m=(110m 的结果)、rate_125m=(125m 处的 rate(queue_depth[10m]))和 deriv_125m=(125m 处的 deriv)。120m 时队列已大部分处理完毕。
值下降时,rate 会认为计数器重启了并加以修正。仪表的下降不是重启,而是真实的变化。
隐藏在一小时平均值里的峰值吞吐量
在 /root/pca-staleness/07-peak.promql 中,用子查询写出最近 1 小时内、每隔 1 分钟观察到的 rate(batch_rows_processed_total[5m]) 的最大值。在 80m 处求值,并在 /root/pca-staleness/07-peak.txt 中写下 peak_rows_per_sec= 和 hour_avg_rows_per_sec=(rate(...[1h]) 的结果)。
식[범위:간격](占位符依次为表达式、范围和间隔)会每隔一个间隔重新对内层表达式求值,生成范围向量。之后就可以在其上使用 *_over_time。
目标消失后仍会持续响起的告警规则
在 /root/pca-staleness/08-rules.yml 中写一个分组和告警 BatchExportStale。expr 使用像第 4 步那样在目标消失后仍有结果的表达式,for: 10m,labels.severity: ticket。用 promtool check rules 检查语法。评分器会用同样的时间序列生成规则测试,确认 60m、125m、135m(pending)时没有触发的告警,而 145m、150m、200m、235m 时 job="nightly-export"、severity="ticket" 的告警处于触发状态。
absent() 会把等号匹配器中的标签附加到结果上。两个分支的标签相同,告警就不会中断而是持续下去。for 是触发前的等待时间。