PCA — Prometheus Certified Associate
Missing measurements and delayed alerts
In one line
A value of 0, having no sample to select right now, and no new sample arriving are different from one another. You also have to distinguish the pending state where an alert first satisfies its condition from the actual firing state. This lab verifies, with real rule tests in virtual time, how that difference changes an alert's wait time. It does not widen the fact that an alert is firing into a claim that a notification email was delivered.
Why this was needed
The fix of "if there is no data, fill it with 0 and the graph gets clean" can make a report look good, but it also risks rewriting a vanished measurement as actually having no requests. A payment service having no requests and an exporter no longer providing the payment counter differ even in how you respond. In the former you analyze demand, and in the latter you check the instrumentation and the collection path.
An alert is also not finished with a comparison of a single number. If the condition is satisfied for a moment and then disappears but the earlier wait time keeps being added up, it can fire too early. Conversely, in a test that leaves out data, you assumed the condition disappeared, but if the real engine selects the previous sample, the alert wait can continue contrary to expectation. Do not be confident about behavior that involves the time axis just from reading an explanation; check it in a test that arranges events before and after.
How it works
This lab has the limited promise that each route must have two input time series. You attach to the basic rate rule a protection condition that keeps only routes whose number of currently selected time series is 2. It is not a recommendation that every service must have exactly two instances. In this lab it is used as a small device that compares the number of current time series with the existence of samples.
In the scenario where an explicit stale marker is inserted, /checkout's b disappears from the current selection result. The count at minute 6 is 1. If you remove the protection condition, a request rate computed from old range samples may keep coming out, but the protected rate has no /checkout sample. The two instances of /status and the value 2 remain as they are. This is not a global empty vector but a shortage of information on one route.
In the real no-request case, on the other hand, the counters of a and b are observed as 0 throughout. The current number of time series is 2 and the protected request rate is 0. If you write the expected samples as an empty list, you wrongly reject this normal state. When you check absence in a test, you must write that the label set is absent from the result, and when you check 0, that the label set exists and its value is 0.
A missing sample and a stale marker are not the same input
In a test's values, _ marks that there is no sample at that position, and stale is a marker that
explicitly states a time series is stale. In the experiment, when only b's sample at minute 6 was left as _, the engine
selected the previous sample from minute 5. The sample age at minute 6 is 60 seconds, the current number of time series is still 2, and the protected rate is 11.
If you change the same position to stale, the protected /checkout result disappears.
This difference arises because selecting the latest value does not necessarily require a new sample at every evaluation time. The previous sample within the lookback can be selected. So the fact that count is 2 is not a guarantee that the current freshness of the two inputs is sufficient or that a real scrape has just succeeded. The identity of the instances also cannot be confirmed from count alone. For a production verdict, you have to design separate promises about the expected target list, the collection status, and the sample age. This unit does not claim to have implemented that promise, and has the student report the limits of the simple count protection condition.
Also, you must not generalize that a real HTTP scrape failure is always handled exactly like _ in a unit test.
A real collector generates stale markers depending on the situation, and target removal and timestamp
settings also have their own separate meanings. Here we feed time series prepared in advance into the rule engine.
Verification of real network collection is distinguished from the independent Prometheus experiment of an earlier unit.
Check the alert wait time on both sides
The alert PcaHighRequestRate is designed to fire only if the protected rate stays above 10.5 for for: 2m. The evaluation interval is 1 minute. The actual results of the normal input are pending at minute 5, pending at minute 6, and firing at minute 7. You have to check both the time the threshold was first crossed and the firing time to detect even a rule that removed for. If you check only the firing at minute 7, a rule that fired earlier is also firing at that time, so a wrong implementation could pass.
If you insert stale at minute 6, in the middle of the alert wait, the condition disappears. Even if the input returns at minute 7, it does not become firing right away and pending starts again. It is also pending at minute 8 and firing at minute 9. Conversely, in the case where only one sample is omitted at minute 6, the condition continues by selection of the previous sample, so it is firing at minute 7. The student test must have both scenarios.
alert_rule_test compares firing alerts at the specified time. An expectation of no firing alone cannot distinguish pending from a completely inactive state. That is why promql_expr_test also checks the alertstate label of ALERTS. The stale case at minute 6 has no ALERTS sample, and the simple missing-sample case has a pending sample. It becomes a much stronger time-axis contract than the single sentence "it has not fired yet."
What it looks like in the field
In alert review, you do not capture just one scene after firing. You test separately just before crossing the threshold, the first crossing, waiting, the firing boundary, and a data gap and its return. Mutations that shorten for to 1 minute or lengthen it to 3 minutes must be detected as firing too early and firing too late respectively. A change that blindly lengthens the wait because the alert is too noisy is also a decision that changes the requirements. How late you can afford to learn of which outage should be discussed together with the team's response budget.
Finally, report what firing means narrowly and precisely. What this test confirms is the state the rule engine generated. Alertmanager's grouping, silencing, and routing, and email delivery are separate stages, and this VM experiment sends no external notifications. The habit of not writing "alerting is fine" for delivery you did not confirm is an operational skill as important as the quality of the tests.
What you will do in the next lab
In coverage-rules.yml, you remove the wrong handling that fills missing values with 0, and you preserve real 0s. In coverage-tests.yml, you fix the wrong expected sample for the missing route. After that, you set the wait time in alert-rules.yml and, in alert-tests.yml, fix the wrong answer that expected firing right after the return. The observation keeps the actual engine output. In the final report, you state the freshness limit of count and the scope of external delivery that this test did not check.