CNPE — Cloud Native Platform Engineer
Repairing three broken boundaries in alert delivery
Goal
On a real k3s and Prometheus Operator, you fix in turn a rule that is not selected, an alert that only fires, and an Alertmanager that does not notify. After actually receiving the current Pod's outage and recovery webhooks, you implement a diagnostic function that distinguishes lack of evidence from contradiction in the evidence.
Why it matters
The Kubernetes API having stored a rule does not mean that Prometheus read that rule. Firing does not mean it was delivered to a person either. Only by fixing each boundary of the official configuration yourself and querying through the API can you tell apart a wrong setting, a state still being applied, and an actual delivery failure.
Operator v0.94.0, Prometheus v3.14.0, Alertmanager v0.34.0, and the issuance app are prepared inside a personal VM. This is a k3s with running processes and is not KWOK. The installation can take several minutes. The work is estimated at 55 minutes. If you need more, extend the time before it expires, and download any records you need before it ends. The VM and files are reclaimed when the session ends. Do not apply anything to a production cluster outside the VM.
What is prepared
- Namespaces platform-monitoring and team-lab. The rule provision-ready, Prometheus/Alertmanager platform
- ServiceMonitor workbench: selects the Service label app=workbench and the Service port name web
- Rule: platform_provision_ready == 0, for 10s, a warning label and description. Keep the expression and spec
- A hypothetical app whose metrics and process can be healthy while its issuance dependency is broken
- Observation tool: python3 /opt/fixtures/cnpe_operator_lab.py
- VM internal ports: app 30088, Prometheus 30090, Alertmanager 30093
observe queries once, and capture followed by a step number waits up to 150 seconds and then saves the specified JSON. Neither fixes the configuration. If it times out, do not reinstall; investigate the current configuration and behavior and then observe again. grade followed by a step number does not modify the saved files.
Steps
- Query the prepared ServiceMonitor and PrometheusRule and save /root/cnpe-alerts/baseline.json with capture 1. Confirm that the metric target is up and that the rule is stored in the API. At first the rule is not selected, so it is not loaded. If you collect again later, you record the actual state at that time, and you do not fabricate the initial failure.
- Fix only the labhub.io/rules label of the team-lab namespace to allowed. Do not widen Prometheus's selector to all namespaces. Save namespace.json with capture 2, and check whether other rule selection conditions remain.
- Attach the label platform=training to the PrometheusRule provision-ready in team-lab. Do not change the rule's spec. Save selected.json with capture 3 and check that the provision rule group was actually loaded in the Prometheus API.
- Send the JSON {"ready":false} to POST /mode of the VM-internal issuance app to create a dependency failure. Save firing.json with capture 4. Confirm that the issuance response is 503 and that Prometheus is firing but there is no alert in Alertmanager or the webhook. Do not inject failures into external services.
- Add spec.alerting.alertmanagers to the Prometheus platform in platform-monitoring. It is a single entry with namespace=platform-monitoring, name=alertmanager, port=web. Keep the prepared query RBAC. Save received.json with capture 5 and separate Alertmanager receipt from no webhook receipt.
- In /root/cnpe-alerts/alertmanager.yaml, write the internal receiver. route.receiver=local-webhook, group_by=[alertname], group_wait=1s, group_interval=5s, repeat_interval=1h. The local-webhook in receivers uses the URL http://workbench.team-lab.svc:8080/hook과 with send_resolved=true. After applying it as the Secret alertmanager-platform in the same namespace with the key alertmanager.yaml, save notified.json with capture 6. Do not send anything outside this URL.
- First confirm the current Pod's firing webhook. Recover by sending {"ready":true} to POST /mode, and save recovered.json with capture 7. Check the 201 response, the clearing of active alerts, and the receipt of both firing and resolved for the same Pod. Do not accept the old Pod's resolved as this recovery.
- Implement diagnose(e) in /root/cnpe-alerts/diagnose.py. Following the diagnostic contract below, it returns a single string. Treat no observation, contradiction, omission, and booleans written as numbers or strings as unknown, and do not lump stored, firing, and received into a single success.
Step 8 diagnostic contract
e needs all seven real bools: observed, selected, loaded, firing, received, notified, and recovered. received and notified are the history of the same outage having passed the respective boundary, and firing is whether it is currently firing. This function receives an already compared summary of the evidence.
- If types or required fields are wrong or observed=False, unknown
- selected=False: if the following five fields are all False, not_selected, otherwise unknown
- selected=True, loaded=False: if the following four fields are all False, not_loaded, otherwise unknown
- With selected and loaded=True and recovered=True: recovered only when firing=False and received and notified=True
- Under the same conditions with recovered=False and notified=True: notified only when firing and received=True
- With notified and recovered=False and received=True: receiver only when firing=True
- With received, notified, and recovered=False: delivery if firing=True, inactive if False
- A combination that fails the "only when" conditions written above is unknown
Reference
A healthy metric scrape, up, is not an indicator of business success. The gauge in this lab is a value that reproduces the dependency state and is not a production SLO. Do not decide the state of an individual alert from the group's status alone. This step's records are compared with the rule UID, the pod UID, and the rule content. If you recreated the Pod, you must observe again. This file comparison is not a cryptographic remote attestation or an anti-cheating device. Steps 1, 4, 5, and 6 check the preserved past observations, and steps 2, 3, and 7 also check the current state. Even after a healthy recovery, do not overwrite the past outage record with a healthy JSON. The receipt records exist only in the app's memory and disappear on restart. Preparing a step requests only the earlier configuration and does not create the current answer or observation files. In step 7, wait for the firing receipt before you recover. Step 8 is an independent coding task.
External webhook, email, and chat delivery, a high-availability alert path, and real-user SLO verification are not covered. Official documentation: Rule selection · ServiceMonitor troubleshooting · Alert states.
Distinguish metrics from the stored rule
Query the prepared ServiceMonitor and PrometheusRule and save /root/cnpe-alerts/baseline.json with capture 1. Confirm that the metric target is up and that the rule is stored in the API. At first the rule is not selected, so it is not loaded. If you collect again later, you record the actual state at that time, and you do not fabricate the initial failure.
Even if targets is up, provision may not be in rules. The two APIs answer different questions.
Fix namespace selection narrowly
Fix only the labhub.io/rules label of the team-lab namespace to allowed. Do not widen Prometheus's selector to all namespaces. Save namespace.json with capture 2, and check whether other rule selection conditions remain.
ruleNamespaceSelector reads the Namespace's labels. The PrometheusRule labels are the next condition.
Fix the label and confirm the actual load
Attach the label platform=training to the PrometheusRule provision-ready in team-lab. Do not change the rule's spec. Save selected.json with capture 3 and check that the provision rule group was actually loaded in the Prometheus API.
It takes time for the actual load to happen right after kubectl succeeds. Do not look only at the declaration file; check /api/v1/rules를 as well.
An outage that fires but is not delivered
Send the JSON {"ready":false} to POST /mode of the VM-internal issuance app to create a dependency failure. Save firing.json with capture 4. Confirm that the issuance response is 503 and that Prometheus is firing but there is no alert in Alertmanager or the webhook. Do not inject failures into external services.
for: 10s requires the condition to persist across real evaluations. A successful /healthz is not a successful /provision.
Connect to Alertmanager and confirm receipt
Add spec.alerting.alertmanagers to the Prometheus platform in platform-monitoring. It is a single entry with namespace=platform-monitoring, name=alertmanager, port=web. Keep the prepared query RBAC. Save received.json with capture 5 and separate Alertmanager receipt from no webhook receipt.
Before fixing the receiver, look at which Alertmanager Prometheus sends to and at the Endpoints read permission.
Actually send through the internal webhook
In /root/cnpe-alerts/alertmanager.yaml, write the internal receiver. route.receiver=local-webhook, group_by=[alertname], group_wait=1s, group_interval=5s, repeat_interval=1h. The local-webhook in receivers uses the URL http://workbench.team-lab.svc:8080/hook과 with send_resolved=true. After applying it as the Secret alertmanager-platform in the same namespace with the key alertmanager.yaml, save notified.json with capture 6. Do not send anything outside this URL.
Not only the Secret name but also the key name and the route's receiver reference must match. Wait for the configuration to be applied too.
Compare whether it is this outage's recovery
First confirm the current Pod's firing webhook. Recover by sending {"ready":true} to POST /mode, and save recovered.json with capture 7. Check the 201 response, the clearing of active alerts, and the receipt of both firing and resolved for the same Pod. Do not accept the old Pod's resolved as this recovery.
Do not look only at the group's single status value; compare the individual status and pod label inside alerts.
A diagnostic tool that does not turn empty evidence into success
Implement diagnose(e) in /root/cnpe-alerts/diagnose.py. Following the diagnostic contract below, it returns a single string. Treat no observation, contradiction, omission, and booleans written as numbers or strings as unknown, and do not lump stored, firing, and received into a single success.
After confirming that all required fields are real bools, look from the earliest boundary. A downstream success together with an upstream failure is a contradiction.