TT Lab
Get started
Learn Learning paths Courses

PCA — Prometheus Certified Associate

Prometheus With Metrics Actually Flowing

Continue in TT Lab

This lab runs on a real Prometheus

Prometheus, Alertmanager, and node-exporter are actually running inside the VM. Because node-exporter emits real metrics, PromQL deals with real data, rules are actually evaluated, and alerts actually fire and reach Alertmanager.

In the fake cluster that the other labs of the PCA course run on, there is no Prometheus itself. All you could do was practice writing queries on paper.

It does not use kube-prometheus-stack. If an Operator writes the configuration for you, you never touch prometheus.yml, and that file is exactly what the PCA asks about.

The first startup takes about 3 minutes.

What is prepared

Prometheus     http://127.0.0.1:30090   설정은 /etc/prometheus/prometheus.yml
Alertmanager   http://127.0.0.1:30093
규칙           /etc/prometheus/rules/rules.yml
대상           node-exporter(데몬셋) · demo-app · Prometheus 자신
질의 도구      promq 'up'

Goal

From the scrape configuration to an alert reaching Alertmanager, you connect the whole path yourself.

Why it matters

The place you get stuck most often when working with Prometheus is not syntax.

The last is especially valuable. You believed "if the service dies, an alert will come," but in a deployment incident where a Pod disappears entirely, no alert comes at all.

Steps

  1. Check what is being scraped right now and record it in /root/pca/targets.txt.
  2. Make it find Pods automatically with kubernetes_sd_configs and record the result in /root/pca/sd.txt.
  3. With relabel_configs, keep only those that have the annotation and create an app label, and record it in /root/pca/relabel.txt.
  4. Write PromQL and record it in /root/pca/promql.txt. Three lines are needed: instant=, range=, and rate_result=.
  5. Create one recording rule (the name follows the level:metric:operation convention) and record in /root/pca/recording.txt that a result actually comes out.
  6. Create an alerting rule, actually break the target so that it reaches firing, and record it in /root/pca/alert.txt. Use a for clause as well.
  7. Make Prometheus send to Alertmanager and record what actually arrived in /root/pca/am.txt.
  8. In /root/pca/report.md, write three lines, up_targets=, alert_fired=yes, and recording_rule=, along with an explanation.

Notes

What is being scraped right now

Check what is being scraped right now and record it in /root/pca/targets.txt.

Look at /api/v1/targets together with the current prometheus.yml.

Do not write targets by hand

Make it find Pods automatically with kubernetes_sd_configs and record the result in /root/pca/sd.txt.

If you use role: pod of kubernetes_sd_configs, the list follows Pods as they come and go.

What to keep and what to drop

With relabel_configs, keep only those that have the annotation and create an app label, and record it in /root/pca/relabel.txt.

Labels that start with __meta_ exist only before relabeling. Filter with keep and create a new label with target_label.

Querying with real data

Write PromQL and record it in /root/pca/promql.txt. Three lines are needed: instant=, range=, and rate_result=.

Three lines are needed: instant=, range=, and rate_result=. Use rate() for counters.

Precomputing

Create one recording rule (the name follows the level:metric:operation convention) and record in /root/pca/recording.txt that a result actually comes out.

Follow the level:metric:operation convention for the name. A colon must be in it to tell it apart from the raw metric.

Actually firing an alert

Create an alerting rule, actually break the target so that it reaches firing, and record it in /root/pca/alert.txt. Use a for clause as well.

up == 0 is true only when the target exists but does not respond. If you delete the Pod, the time series itself disappears.

There has to be somewhere for the alert to go

Make Prometheus send to Alertmanager and record what actually arrived in /root/pca/am.txt.

If you do not write alerting.alertmanagers, even when a rule fires it goes nowhere.

What you learned

In /root/pca/report.md, write three lines, up_targets=, alert_fired=yes, and recording_rule=, along with an explanation.

Write three lines, up_targets=, alert_fired=yes, and recording_rule=, along with an explanation.