CNPE — Cloud Native Platform Engineer
Distinguishing firing alerts from delivered notifications
One-line summary
Prometheus firing, Alertmanager receipt, webhook arrival, and business recovery are different pieces of evidence. Even if the alert name is the same, you must not count another Pod's recovery as the recovery of this outage.
Why this was needed
You finally loaded the issuance platform's rule. This time Prometheus shows firing. The operator closed the report saying "the alerts are fine now," but no request arrived at the notification server. This was because there was no configuration for Prometheus to send to Alertmanager. When we fixed the connection, this time the alert appeared in the Alertmanager API. Even so, nothing arrived at the webhook. The selected receiver was an empty receiver configuration with no actual delivery target.
The habit of approving the later boundaries on the grounds that an earlier boundary succeeded creates the same problem in deployment and in alerting. This time you do not need to make the outage bigger or lower the threshold. First investigate between the place where the alert was last confirmed and the place where it first disappeared.
How it works
If a rule's expression produces a result for some label set, the alert for that target becomes active. If you specified for, it waits to see whether the condition persists at each evaluation and then moves to firing. So for: 10s does not simply mean measuring 10 seconds after the install command. When the first valid evaluation happened also has an effect. For the relationship between pending and firing, see the Prometheus rules documentation.
The connection from Prometheus to Alertmanager and the notification from Alertmanager to the actual recipient are configured separately. Alertmanager is responsible for grouping, inhibition, silencing, and sending. This distinction appears in the official alerting overview. In the lab, instead of an external messenger, you use an HTTP webhook inside the personal VM to check whether a real request came in. You look at the content of the request received, not at the tool's settings screen or a success message.
Written as questions, the diagnostic order is as follows.
| What you last confirmed | What is not there yet | The boundary to investigate first |
|---|---|---|
| The rule is stored | The rule group in the running instance | Selection scope and configuration reflection |
| Prometheus firing | The corresponding target in Alertmanager | Connection, discovery permissions, access path |
| Alertmanager receipt | The webhook request | route, receiver, delivery failure |
| Webhook firing arrival | The recovery of the same target and business success | Dependency recovery, a new evaluation, recovery delivery |
If there is no recipient, an empty receiver can also be a valid configuration. You must test separately whether a syntactically correct configuration performs the delivery you want for the business. When you use a Secret, the name and the inner key must match the contract the consumer expects. This lab's configuration has no secret values, but if you actually put in an API key or a password, do not copy it into source, screenshots, or observation records.
What it looks like in the field
In the preparatory experiment, we replaced the app with a new Pod and then caused the same outage again. Within one webhook group came the new Pod's firing and the old Pod's resolved together. The alert name and the lab label were the same. The judgment "if there is even one resolved, it is recovered" is wrong at this moment. You had to compare each alerts entry against the current Pod's labels.
Also distinguish the group's single status value from the status of the individual alerts within the group. When several targets are bundled into one group, a single group status cannot represent the state of all the targets. In a real environment, which label to take as the event's identifier, and whether to link or separate events before and after a restart, must also be decided as a contract. This task puts the rule UID and the pod UID into the observation file so that records from a different run are not reused by mistake. This is not an electronic signature proving the authenticity of the file.
Give names to observation gaps too. If you interpret an empty object after an API error as normal, the false conclusion "there is no outage" is produced. If the input is missing or contradictory, return unknown and first redo the needed queries. This is also the reason not to treat the number 0 and a real boolean False as the same. You keep a wrong data format from looking like a state that happens to be correct.
A past outage record and the current state serve different purposes. After recovery, the current API may have no active alerts, but the incident report must still keep the record of what fired and was received earlier. If you overwrite the saved JSON entirely with the current state, the material to explain the cause disappears. Conversely, you must not say the current service is healthy holding only an old healthy JSON. The final recovery step rechecks not only the preserved records but also the current business response and the active alert state.
What to do in the next lab
You break the issuance dependency and leave separate records for the point where there is only Prometheus firing, the point where it reached only Alertmanager, and the webhook receipt and recovery points. At the end, you build a diagnostic tool that classifies the state with a strict boolean contract. Rather than feeding in just one success case, you feed in omissions, contradictions, and wrong types and confirm that they are rejected. External notification delivery, multi-node high availability, and meeting a production SLO are outside the scope this lab proves.