PCA — Prometheus Certified Associate
Prometheus With Metrics Actually Flowing
This lab runs on a real Prometheus
Prometheus, Alertmanager, and node-exporter are actually running inside the VM. Because node-exporter emits real metrics, PromQL deals with real data, rules are actually evaluated, and alerts actually fire and reach Alertmanager.
In the fake cluster that the other labs of the PCA course run on, there is no Prometheus itself. All you could do was practice writing queries on paper.
It does not use kube-prometheus-stack. If an Operator writes the configuration for you,
you never touch prometheus.yml, and that file is exactly what the PCA asks about.
The first startup takes about 3 minutes.
What is prepared
Prometheus http://127.0.0.1:30090 설정은 /etc/prometheus/prometheus.yml
Alertmanager http://127.0.0.1:30093
규칙 /etc/prometheus/rules/rules.yml
대상 node-exporter(데몬셋) · demo-app · Prometheus 자신
질의 도구 promq 'up'
Goal
From the scrape configuration to an alert reaching Alertmanager, you connect the whole path yourself.
Why it matters
The place you get stuck most often when working with Prometheus is not syntax.
- You fixed the configuration but it does not take effect. A ConfigMap volume is synchronized periodically by the kubelet, so it does not change immediately (about 60 seconds). If you hit
reloadbefore propagation, it rereads the old configuration, andreloadreturns 200. - The alert does not arrive. It is common that the rule fired but
alerting.alertmanagerswas not written, so there is nowhere to send it. - The
up == 0alert does not trigger. When a target disappears, theuptime series itself goes away.up == 0is true only when the target exists but does not respond.
The last is especially valuable. You believed "if the service dies, an alert will come," but in a deployment incident where a Pod disappears entirely, no alert comes at all.
Steps
- Check what is being scraped right now and record it in
/root/pca/targets.txt. - Make it find Pods automatically with
kubernetes_sd_configsand record the result in/root/pca/sd.txt. - With
relabel_configs, keep only those that have the annotation and create anapplabel, and record it in/root/pca/relabel.txt. - Write PromQL and record it in
/root/pca/promql.txt. Three lines are needed:instant=,range=, andrate_result=. - Create one recording rule (the name follows the
level:metric:operationconvention) and record in/root/pca/recording.txtthat a result actually comes out. - Create an alerting rule, actually break the target so that it reaches
firing, and record it in/root/pca/alert.txt. Use aforclause as well. - Make Prometheus send to Alertmanager and record what actually arrived in
/root/pca/am.txt. - In
/root/pca/report.md, write three lines,up_targets=,alert_fired=yes, andrecording_rule=, along with an explanation.
Notes
- Query with
promq 'up'. The result is JSON, so filter it with| jq. - After you fix the configuration file, run
curl -XPOST http://127.0.0.1:30090/-/reload. The Pod sees that directory directly as a hostPath, so it takes effect the moment you fix it. - If reload is not 200, the configuration has a syntax error. In that case the old configuration stays alive, so the service does not die.
- Check the target status with
curl -s $P/api/v1/targets | jq.droppedTargetsholds what was filtered out by relabeling. - To create
up == 0in step 6, the target must exist and not respond. The most reliable way is to pointstatic_configsat a port nobody is listening on. - Common mistake 1: trying to create
up == 0by deleting the Pod. The target disappears from the list and the time series itself goes away. That case has to be caught withabsent(). - Common mistake 2: not using a colon in a recording rule name. You can no longer tell the metric a rule created from the raw metric by name alone.
What is being scraped right now
Check what is being scraped right now and record it in /root/pca/targets.txt.
Look at /api/v1/targets together with the current prometheus.yml.
Do not write targets by hand
Make it find Pods automatically with kubernetes_sd_configs and record the result in /root/pca/sd.txt.
If you use role: pod of kubernetes_sd_configs, the list follows Pods as they come and go.
What to keep and what to drop
With relabel_configs, keep only those that have the annotation and create an app label, and record it in /root/pca/relabel.txt.
Labels that start with __meta_ exist only before relabeling. Filter with keep and create a new label with target_label.
Querying with real data
Write PromQL and record it in /root/pca/promql.txt. Three lines are needed: instant=, range=, and rate_result=.
Three lines are needed: instant=, range=, and rate_result=. Use rate() for counters.
Precomputing
Create one recording rule (the name follows the level:metric:operation convention) and record in /root/pca/recording.txt that a result actually comes out.
Follow the level:metric:operation convention for the name. A colon must be in it to tell it apart from the raw metric.
Actually firing an alert
Create an alerting rule, actually break the target so that it reaches firing, and record it in /root/pca/alert.txt. Use a for clause as well.
up == 0 is true only when the target exists but does not respond. If you delete the Pod, the time series itself disappears.
There has to be somewhere for the alert to go
Make Prometheus send to Alertmanager and record what actually arrived in /root/pca/am.txt.
If you do not write alerting.alertmanagers, even when a rule fires it goes nowhere.
What you learned
In /root/pca/report.md, write three lines, up_targets=, alert_fired=yes, and recording_rule=, along with an explanation.
Write three lines, up_targets=, alert_fired=yes, and recording_rule=, along with an explanation.