PCA — Prometheus Certified Associate
Recording Rules, Alerting Rules, and Testing Them
Goal
You design recording rules in a two-level hierarchy, put alerting rules on top of them, and even write a unit test with calculated expected values. Finally, you move the same content into the CR form that the Prometheus Operator reads.
Why it matters
A recording rule is not "a cache that makes queries fast" but a design decision that moves the time of computation. The load at query time moves to evaluation time, and in return new time series exist permanently. That is why you have to multiply out the count before creating one. With 120 routes and 5 window types, one rule is 600 time series, and 50 rules is 30,000. The reason for splitting into layers is the same. Once layer 1 has scanned the raw data once, layer 2 and the alerting rules read only that result, so the raw scan is not repeated as many times as there are rules. And rules are production code. Production code is not deployed without tests.
Steps
- Create
/root/pca-rules/recording.ymland write the first group undergroupsasname: http_sli,interval: 30s. - In that group's
rules, put two layer-1 rules.record: route:http_requests:rate5missum(rate(http_requests_total[5m])) by (route), andrecord: route:http_requests_errors:rate5missum(rate(http_requests_total{status_class="5xx"}[5m])) by (route). - Add the layer-2 rule
record: route:http_error_ratio:rate5mbelow layer 1. The expression isroute:http_requests_errors:rate5mdivided byroute:http_requests:rate5m, and you must not use the rawhttp_requests_totalagain. - Add
record: route:http_request_duration_seconds:p99_rate5m. The expression ishistogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route)). The recording rules in this file total 4. - Create
/root/pca-rules/alerting.ymland write the alertCheckoutHighErrorRatiounder the groupname: http_alerts.exprisroute:http_error_ratio:rate5m > 0.01,for: 5m,labels.severity: page, andannotationshassummaryandrunbook_url(a URL starting with http). - In
/root/pca-rules/recording_test.yml, write a promtool unit test.rule_filesisrecording.yml,evaluation_interval: 30s,tests[0].interval: 15s,input_serieshas two time series, 2xx (0+150x40) and 5xx (0+3x40), andpromql_expr_testhasexpr: route:http_error_ratio:rate5m,eval_time: 8m, and thevalueofexp_samplesis 3 divided by 153 (starting with 0.0196). - Create the namespace
pca-rules, and in/root/pca-rules/prometheusrule.yamlwrite a PrometheusRule.apiVersion: monitoring.coreos.com/v1,kind: PrometheusRule,metadata.name: checkout-sli,metadata.namespace: pca-rules,metadata.labels.release: kube-prometheus-stack,spec.groups[0].name: checkout_sli, and put bothrecord: route:http_error_ratio:rate5mandalert: CheckoutHighErrorRatioinside it.
Notes
exprmay be written on one line or across several lines with anexpr: |block. Grading ignores whitespace.- Calculate the expected value in step 6 yourself. If 2xx grows by 150 and 5xx by 3 every 15 seconds, the rate ratio is 3 / (150 + 3).
- Common mistake 1: writing layer 2 above layer 1. A group is evaluated from top to bottom.
- Common mistake 2: leaving the colon out of a recording rule name. The colon is the mark of a derived time series.
- Common mistake 3: leaving the
releaselabel off the PrometheusRule. The Operator's ruleSelector picks by this label.
Creating the rule group shell
Create /root/pca-rules/recording.yml and write the first group under groups as name: http_sli, interval: 30s.
The top level of a rule file is a groups list. Each group has a name and an optional interval, and if you omit interval, global.evaluation_interval is used. The rules in the same group are evaluated in order.
Layer 1 — scan the raw data only once
In that group's rules, put two layer-1 rules. record: route:http_requests:rate5m is sum(rate(http_requests_total[5m])) by (route), and record: route:http_requests_errors:rate5m is sum(rate(http_requests_total{status_class="5xx"}[5m])) by (route).
Write the new time series name in the record field and the expression in expr. Keep the order that puts rate inside sum, and both rules must aggregate along the same dimension (route) so that you can divide them later.
Layer 2 — reference only layer 1
Add the layer-2 rule record: route:http_error_ratio:rate5m below layer 1. The expression is route:http_requests_errors:rate5m divided by route:http_requests:rate5m, and you must not use the raw http_requests_total again.
If the ratio rule scans the raw metric again, splitting into layers is pointless. Divide using only the two time series names you made earlier. Within a group, rules are evaluated from top to bottom, so the order matters too.
p99 recording rule
Add record: route:http_request_duration_seconds:p99_rate5m. The expression is histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route)). The recording rules in this file total 4.
The name follows the level:metric:operations convention. A name without a colon cannot be told apart from a raw metric. In the expression, apply rate to the buckets and be sure to keep le in the aggregation.
Alerting rule file
Create /root/pca-rules/alerting.yml and write the alert CheckoutHighErrorRatio under the group name: http_alerts. expr is route:http_error_ratio:rate5m > 0.01, for: 5m, labels.severity: page, and annotations has summary and runbook_url (a URL starting with http).
An alerting rule uses the alert field instead of record. The expression should reference the recording rule result so that it does not sweep the raw data at every evaluation. for is the time the condition is continuously true, and if it is false even once in between, the timer goes back to 0.
promtool unit test file
In /root/pca-rules/recording_test.yml, write a promtool unit test. rule_files is recording.yml, evaluation_interval: 30s, tests[0].interval: 15s, input_series has two time series, 2xx (0+150x40) and 5xx (0+3x40), and promql_expr_test has expr: route:http_error_ratio:rate5m, eval_time: 8m, and the value of exp_samples is 3 divided by 153 (starting with 0.0196).
The values of input_series use the 시작+증가x횟수 syntax (the placeholders are the start, the increment, and the count). If 2xx grows by 150 and 5xx by 3 every 15 seconds, the error ratio is 3 divided by 153. If you calculate the expected value yourself and write it down, CI catches it when the rule's meaning changes later.
Moving into a PrometheusRule CR
Create the namespace pca-rules, and in /root/pca-rules/prometheusrule.yaml write a PrometheusRule. apiVersion: monitoring.coreos.com/v1, kind: PrometheusRule, metadata.name: checkout-sli, metadata.namespace: pca-rules, metadata.labels.release: kube-prometheus-stack, spec.groups[0].name: checkout_sli, and put both record: route:http_error_ratio:rate5m and alert: CheckoutHighErrorRatio inside it.
The Operator picks CRs with ruleSelector, so if the labels do not match, the file is ignored even though it exists. The structure of spec.groups is the same as a rule file, and one CR can hold both recording rules and alerting rules. This environment has no CRD, so you only write the file and actually create only the namespace.