PCA — Prometheus Certified Associate
Counter resets hidden by healthy test data
In one line
That a rule file's syntax is correct and that the rule produces the right numbers are different claims. If you combine counters first, the resets of individual processes can be hidden. A test that has only normal data also lets that very wrong rule pass. In this unit, you put counterexamples into a real PromQL engine to tell the two claims apart, and you fix the recording rule and the test file yourself.
Why this was needed
Think of the day you fixed a payment service's request rate dashboard. In review, promtool check rules succeeded, and the graph of usual traffic was smooth too. But only in the period when one instance restarted, the dashboard's request rate dropped. It was not a syntax error but an error in the order of calculation, so the parser could not find it. Since the aggregated numbers look plausible and the alerts are quiet, people tend to feel reassured instead.
A test is not a ritual of running a command with test in its name once. It is a device that states what labels and values are expected for which input, and rejects implementations that break that expectation. If the green light comes on even when you put in a wrong implementation, it means the test is missing a situation. You do not need to add tests only after discovering every outage in production. You can first find the information that rules easily hide and model it as a small time series.
How it works
The lab input has two labels, route and instance. For /checkout, the cumulative counter of a grows by 60 every minute and b grows by 600 every minute. So over a long enough window, a is 1 per second and b is 10 per second. For /status there are two instances, each growing by 60 per minute. Only with this contrast path can a test reject an aggregation that also erased route. If you put in only /checkout, it is hard to notice a result with the label wrongly removed just by looking at the numbers.
The correct recording rule computes rate for each raw time series and then sums by route. The comparison rule first records a time series that adds up the raw counters by route and then applies rate to that total. At minute 6 of the normal input, both rules return 11 for /checkout. If you test only this far, you have no grounds to judge which implementation is wrong.
In the counterexample, a goes 0, 60, 120, 180, 0, 60, 120. b keeps increasing. Even at the moment a resets, b's increase is larger, so the total of the two values does not decrease. A rate that can see the raw series accounts for a's reset, but a rate that has already received only the total cannot recover that fact. The key point is that information lost in the summation does not come back later by changing the function name.
The actual Prometheus 3.14.0 results with a fixed 1-minute input and evaluation interval, evaluation at minute 6, and a 5-minute window are a correct /checkout request rate of 10.75 and a wrong request rate of 10. The per-route sum of the raw resets is 1, while the resets of the recorded total is 0. /status is still 2. If you check these four results together, a fix that simply returns the constant 10.75 or removes the route is not accepted either.
The 10.75 here does not mean every request stamped in an event ledger was counted per second. rate uses the samples in the range and boundary extrapolation. If you change the evaluation time, the window, or the input interval, the result can change too. That is why the test file includes not only the input values but also interval, evaluation_interval, and eval_time. If you omit the time conditions and memorize only the numbers, you cannot explain the next case.
The order for reading a test file
First look at what rule_files reads. The lab helper copies the student's step rules to rules.json in a temporary folder. The extension is JSON, but it is valid as the YAML representation that Prometheus accepts. Next you read the labels and values of the input time series, the evaluation interval, the check time, and the expected samples. You may change the order of arrays, but leaving out a needed time series or expected sample means creating a different test.
Each promql_expr_test has expr, eval_time, and exp_samples. In the labels of exp_samples you write who owns the number, and in value you write the expected value. A check where only the numbers match is not enough. A state where the route label has vanished but only the total is the same can also be a wrong result. If you lower the reset test's answer to 10, the wrong rule passes, but that is not fixing the implementation; it is changing the requirement. The input and expected-value contract of this exercise rejects that kind of relaxation.
The lab does not reject an expression that produces the same result just because the string differs. For example, when the only labels of this input are route and instance, sum without(instance) can produce the same result as sum by(route). That does not mean the two expressions are the same on every production input that has other labels added. The scope a test guarantees is the scope of the inputs that went into the test.
What it looks like in the field
In a deployment pipeline, you have both a syntax check and a semantic test. The syntax check quickly flags a wrong function call or a bad field or expression, and the semantic test checks the relationship among labels, values, and time that the service promised. In a review, do not attach just one normal input; attach cases that can shake the calculation, such as a restart, route separation, no requests, and data gaps. Look not only at the result of the changed rule but also at whether existing important cases are maintained. This is because there can be a regression where you fixed the number but the original route separation broke.
Also try intentionally putting in a wrong implementation once. In this experiment, a test that kept only normal data and an empty tests array actually let a wrong aggregation rule pass. A tool's successful exit means there was no problem in the checks it was asked to run, not a guarantee that all the needed checks were done. The habit of reviewing what was checked applies just as it is to deployment gates in general.
What you will learn next
In the next lesson, you first read the time axis of gaps and alerts. Then, in the lab after that, you read the output where the same wrong rule passes the normal test and fails the reset test. After that, you fix the recording rule in rate-rules.yml and correct the wrong expected value in reset-tests.yml. The helper runs the real promtool on a temporary copy of the student files and keeps the syntax and semantic results separately. It does not change the host or production monitoring, and the input is a synthetic time series for learning.