TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Valid syntax, incorrect alerts

Continue in TT Lab

Goal

You fix recording rules and alerting rules yourself, and write tests for normal, reset, gap, and wait-time cases to detect semantic errors.

Why it matters

Syntax success is not semantic success. A wrong rule can pass with only normal data or an empty test. This lab uses the real promtool of a fixed Prometheus 3.14.0 and synthetic time series. It does not change production LabHub or any external Prometheus's configuration, real scrape targets, or Alertmanager. It is a 55-minute lab. Extend it before it expires if needed. When the session ends, the VM and student files are reclaimed.

Prepared files and commands

The working folder is /root/pca-rule-lab. All file names in the steps below are relative to this folder. materials/inputs.json is the fixed input, and materials/expected.json is the full expected result. materials/starter- (the placeholder is the file name) is a starting point with intentional errors. Copy that file to the same file name in the working folder and edit it. Do not overwrite the installed material files themselves. The three rule files are kept separate, so a fix in a later step does not erase an earlier step's answer. Do not guess the answer from the materials and examples; compare against the contract in the text.

python3 /opt/fixtures/pca_rule_lab.py complete N checks the current answer and saves the syntax and semantic check results and the input hash in observation-N.json. If it fails, compare exp and got in the output, fix the answer, and run it again. grade N reruns the real engine but does not change answers or observations. prepare N fills in only the earlier steps. solve N is viewing the answer; it creates only the answers that do not exist and does not overwrite existing partial answers. answer N only prints the answer content to the terminal. You can compare a partial answer with that content and fix it yourself. The test file's rule_files keeps the single rules.json. The helper moves that step's rules into the rules.json of a temporary copy and then runs promtool check rules and promtool test rules. You do not need to create rules.json in the student folder. JSON is valid as a representation of YAML, and you may also write it directly in YAML. The inputs, queries, times, and expected values are a fixed exercise contract, and the order of tests, the order of inputs, the order of samples, and whitespace may differ. Rule expressions are checked by their results. Keep one group, three recording rules and one alert, and the names, order, intervals, and labels as in the starter. Extra rules, templates, external files, and YAML aliases are outside the scope of this exercise. A file must be UTF-8 and at most 64KiB.

Steps

  1. In diagnosis.json, write the booleans syntax_proves_semantics=false and normal_data_sufficient=false and the numbers reset_good_rate=10.75 and reset_bad_rate=10, and run complete 1. In observation-1.json you compare, for the same wrong aggregation rule, syntax success, normal data success, and reset counterexample failure.
  2. Copy materials/starter-rate-rules.yml to rate-rules.yml and fix the expr of pca:requests:rate5m. With rate per raw series and then summing by route, normal minute 6 must give /checkout=11 and /status=2, and reset minute 6 must give /checkout=10.75. Keep the other rules, the group, the 1m interval, and the labels, and run complete 2.
  3. Copy materials/starter-reset-tests.yml to reset-tests.yml and change the expected value of the minute-6 request rate in masked-counter-reset from 10 to 10.75. The inputs, queries, labels, and remaining expected values of the two scenarios are compared against materials/expected.json. complete 3 checks both the success of the student rule and the failure of the wrong answers of aggregation order and route loss.
  4. Copy materials/starter-coverage-rules.yml to coverage-rules.yml. In the expr of pca:requests:ready_rate5m, remove the part that fills missing routes with 0, and keep only routes where the current count by(route) is 2. With an explicit stale there must be no /checkout result, and with real no requests there must be the value 0. Preserve /status=2 and run complete 4.
  5. Copy materials/starter-coverage-tests.yml to coverage-tests.yml. In the expected samples for ready_rate5m at minute 6 in missing-is-not-zero, remove /checkout=0. Keep the 0 sample of real-zero. Also keep missing-sample-is-not-staleness's current count 2, sample age 60 seconds, and request rate 11, and run complete 5. Do not interpret the current number of time series as a freshness guarantee.
  6. Copy materials/starter-alert-rules.yml to alert-rules.yml and set for: 2m on PcaHighRequestRate. Keep expr as the protected rate > 10.5 and severity as practice. Normal data must be pending at minutes 5 and 6 and firing at minute 7, and after a stale at minute 6 it must be pending at minutes 7 and 8 and firing at minute 9. Run complete 6 to check all six scenarios.
  7. Copy materials/starter-alert-tests.yml to alert-tests.yml. Fix the expected firing alert at minute 7 of gap-restarts-pending to an empty exp_alerts. Keep the firing at minute 9, the pending, absent, and firing checks of ALERTS, and the firing at minute 7 of the simple missing sample. complete 7 also checks the four wrong answers: no for, 1 minute, 3 minutes, and no protection condition.
  8. In report.json, write the booleans missing_is_zero=false, count_proves_freshness=false, notification_delivery_tested=false, and production_scraping_tested=false and the strings stale_gap_fires_at="9m" and missing_sample_fires_at="7m". With complete 8, re-verify the rules, the test contract, and the earlier steps' observations. Do not write that you confirmed real notification delivery or production scrapes.

Notes and limitations

These are the results of a synthetic input starting from an initial 0 and a 1-minute evaluation interval. They include the boundary extrapolation of the 5m window, so do not change the input and expect the same numbers. While only a of /checkout resets, b keeps increasing, and /status is an independent contrast path. An explicit stale and a simple missing sample are different inputs. count=2 is not a guarantee of sample freshness or instance identity. The wait time is virtual evaluation time, and you do not actually wait 9 minutes. pending and firing are also checked through the ALERTS labels. Confirming that an alert is firing is not confirming delivery of an external notification. The network behavior of a real scrape is not in this test. The material's hash is a device for finding accidental overwrites and mixing up observations from other labs. It is not a security guarantee that prevents the root of the same VM from tampering with all the code. Grading has a 60-second budget and preparation a 90-second budget, and the engine is limited to an even shorter time. Official rule tests

Dividing responsibility between syntax checks and semantic checks

In diagnosis.json, write the booleans syntax_proves_semantics=false and normal_data_sufficient=false and the numbers reset_good_rate=10.75 and reset_bad_rate=10, and run complete 1. In observation-1.json you compare, for the same wrong aggregation rule, syntax success, normal data success, and reset counterexample failure.

See what result the same rule produces on different inputs.

Writing a recording rule that does not hide resets

Copy materials/starter-rate-rules.yml to rate-rules.yml and fix the expr of pca:requests:rate5m. With rate per raw series and then summing by route, normal minute 6 must give /checkout=11 and /status=2, and reset minute 6 must give /checkout=10.75. Keep the other rules, the group, the 1m interval, and the labels, and run complete 2.

If you apply rate to the total, you can lose the raw resets.

Writing a test that rejects a wrong aggregation

Copy materials/starter-reset-tests.yml to reset-tests.yml and change the expected value of the minute-6 request rate in masked-counter-reset from 10 to 10.75. The inputs, queries, labels, and remaining expected values of the two scenarios are compared against materials/expected.json. complete 3 checks both the success of the student rule and the failure of the wrong answers of aggregation order and route loss.

If you lower the expected value to fit the implementation, the test allows the wrong rule.

Writing a rule that distinguishes missing measurement from a real 0

Copy materials/starter-coverage-rules.yml to coverage-rules.yml. In the expr of pca:requests:ready_rate5m, remove the part that fills missing routes with 0, and keep only routes where the current count by(route) is 2. With an explicit stale there must be no /checkout result, and with real no requests there must be the value 0. Preserve /status=2 and run complete 4.

No current sample and an existing value of 0 are different states.

Leaving gaps and the freshness limit in the test

Copy materials/starter-coverage-tests.yml to coverage-tests.yml. In the expected samples for ready_rate5m at minute 6 in missing-is-not-zero, remove /checkout=0. Keep the 0 sample of real-zero. Also keep missing-sample-is-not-staleness's current count 2, sample age 60 seconds, and request rate 11, and run complete 5. Do not interpret the current number of time series as a freshness guarantee.

Use the timestamp to see whether the previous sample is selected within the lookback.

Setting the time boundaries before and after an alert fires

Copy materials/starter-alert-rules.yml to alert-rules.yml and set for: 2m on PcaHighRequestRate. Keep expr as the protected rate > 10.5 and severity as practice. Normal data must be pending at minutes 5 and 6 and firing at minute 7, and after a stale at minute 6 it must be pending at minutes 7 and 8 and firing at minute 9. Run complete 6 to check all six scenarios.

You also have to check the state just before firing to find a firing that is too early.

Testing whether an interrupted wait restarts

Copy materials/starter-alert-tests.yml to alert-tests.yml. Fix the expected firing alert at minute 7 of gap-restarts-pending to an empty exp_alerts. Keep the firing at minute 9, the pending, absent, and firing checks of ALERTS, and the firing at minute 7 of the simple missing sample. complete 7 also checks the four wrong answers: no for, 1 minute, 3 minutes, and no protection condition.

After the condition is cut off by a stale, for counts again from the time of the return.

Re-verifying the rules and tests together and reporting the scope

In report.json, write the booleans missing_is_zero=false, count_proves_freshness=false, notification_delivery_tested=false, and production_scraping_tested=false and the strings stale_gap_fires_at="9m" and missing_sample_fires_at="7m". With complete 8, re-verify the rules, the test contract, and the earlier steps' observations. Do not write that you confirmed real notification delivery or production scrapes.

Report separately the state confirmed in the test and the behavior of external systems.