CNPE — Cloud Native Platform Engineer
From a saved rule to an evaluated rule
One-line summary
Having created a PrometheusRule means you stored the rule in Kubernetes. Whether that rule is selected, loaded into Prometheus, and evaluated against real time series must each be checked.
Why this was needed
A hypothetical issuance platform team added an outage alert rule. The terminal printed created, and code review also confirmed that the expression was right. A few days later the issuance API returned 503, yet no call came. Searching for the rule name showed it in Kubernetes, but it was not in Prometheus's rules API. The team said "let's make the alert expression more sensitive," but changing the threshold of an expression that is not even being read does not fix this failure.
The problem lay in calling two different successes the same thing. The API server is responsible for storing the object, while the Operator picks the objects to put into the Prometheus it manages and builds the configuration. Even if an application team has permission to store rules, if the design is that every Prometheus reads those rules, the rules of multiple teams get mixed indiscriminately. The selection scope is not an inconvenient barrier but a device that expresses ownership and operational scope.
How it works
| Observation target | Question to check | What it cannot yet prove |
|---|---|---|
| PrometheusRule in Kubernetes | Is the rule object stored? | Operator selection and actual evaluation |
| Labels of the Namespace and the rule | Does it fall within the specified selection scope? | Whether the configuration was reflected in the real process |
| Prometheus rules API | Is the group loaded into the running instance? | Whether the condition is met and whether the alert is delivered |
| targets API | Which addresses are scraped, with what result? | Whether that business function is succeeding |
Selection narrows twice. First, ruleNamespaceSelector picks the namespaces in which to look for rules. Within that scope, ruleSelector checks the PrometheusRule's labels. The meaning of defaults and empty selectors can differ per field, so do not guess "I wrote nothing, so it must select everything." Check the exact behavior in the Operator rule selection guide.
For example, suppose there is an instance that reads only rules with owner=payments among the namespaces with monitoring=yes. Even if you put owner=payments on the rule, if the namespace is outside the scope, it does not read it. If you fixed only the namespace but the rule is owner=orders, it still does not read it. If you change both to select everything at once, it may show up right away, but it can also read other teams' rules and create duplicate alerts or increased cost. Fix only the one condition that is needed and observe again.
Metric scraping is also a separate selection path. A ServiceMonitor finds scrape targets through the Service's labels and port name, and Prometheus selects the ServiceMonitor's namespace and labels. Writing the port number as a string is not the same as using the Service's port name. Insufficient permissions can also appear as a symptom of targets not being visible, like a label mismatch, so following the official troubleshooting order, you check in turn the selected objects, the generated configuration, and the service and discovery permissions.
What it looks like in the field
In this unit's preparatory experiment, the target metric was up but the rule groups were empty. Even after fixing only the namespace label and watching for 75 seconds, the rule did not appear. When we matched the rule labels, it loaded after about 80 seconds. This is a single observed value and not a recommendation to fix the wait time. It is a case that shows that the API change, the projection of the configuration file, the reload, and the evaluation happen at different times, so you should not treat the terminal's configured as the final evidence.
The shape of the failure response matters too. The exporter we first built emitted a correct metric body, but Content-Type was empty, so Prometheus 3's scrape failed. It succeeded in finding the target and failed in reading the content. Changing the labels again does not fix this problem. We read the target's lastError and fixed the producer's HTTP header. You can check this difference against the Prometheus 3 migration guide.
A successful scrape, up, is the observation that the application returned metrics. Even if an external dependency of the issuance function breaks, the exporter process can keep responding normally. Conversely, the absence of a metric does not mean the business function necessarily failed. Record the business state and the observation state separately, and if you do not know one of them, leave that blank as it is.
What to do in the next lesson
In the next reading, you look at the boundary where notifications do not arrive even after the rule is loaded. In the lab after that, without changing the expression of the prepared rule, you fix the namespace label and the rule label separately. You look at the stored object and the actual rules and targets APIs together. After fixing one boundary, record what changed and what stayed the same. The lab uses a real k3s on a personal VM and is not a procedure for applying to a production cluster.