TT Lab
Get started
Learn Learning paths Courses

Observability

Throwing Queries at a Real Prometheus

Continue in TT Lab

Goal

You throw queries at a real Prometheus server and read the numbers that come back. By the end of this lab, you will have written rate, ratio, quantile, and forecast queries yourself and will know why they were wrong when they are wrong.

Why it matters

You can memorize PromQL syntax in a day. The hard part is judging whether the number that came back is right. A histogram_quantile that omits by (le) does not raise an error — it just gives you a wrong number. If you put that on a dashboard, nobody knows for months.

That is why this environment comes preloaded with 12 hours of time series. There are two periods in which errors spike, one period in which only p99 spikes, and a disk that decreases monotonically. You can check the answer by whether the queries you wrote find those events.

Environment

promq "<PromQL>"      쿼리를 던진다. promq -r 로 원본 JSON
lab-status            무엇이 떠 있는지, 대상이 몇 개인지
tail -40 /var/log/lab-init.log     준비 과정 로그

Prometheus is at 127.0.0.1:9090, the sample service exporter is at 9101, and node_exporter is at 9100. Start Grafana with lab-start-grafana and view it through web preview port 3000.

Steps

  1. Find which labels http_requests_total has and write them to /root/obs/01-labels.txt
  2. Write the per-second request rate query to /root/obs/02-rate.promql
  3. Write the 5xx ratio query to /root/obs/03-errrate.promql
  4. Write the p95 latency query to /root/obs/04-p95.promql
  5. Write the five metrics with the most time series to /root/obs/05-top.txt as 이름 시계열수 (name and series count)
  6. Write the recording rules to /etc/prometheus/rules/labhub.yml
  7. Apply the configuration and confirm that the new time series appears
  8. Write the disk exhaustion forecast query to /root/obs/08-predict.promql

Notes

First look at what time series exist

Find which labels http_requests_total has and write them to /root/obs/01-labels.txt

Before writing a query, look at what exists. Scan the metric names with promq 'count by (__name__)({__name__=~".+"})' and look at the labels with promq 'http_requests_total'. Write the label names of http_requests_total to /root/obs/01-labels.txt, one per line (names only, not values).

Per-second request rate

Write the per-second request rate query to /root/obs/02-rate.promql

Write the query to /root/obs/02-rate.promql and check it directly with promq "$(cat /root/obs/02-rate.promql)". The scrape interval is 15 seconds, so the range must be at least four times that so there is no moment when only one sample is captured. rate must be inside sum — if you put it outside, counter resets are handled wrongly.

5xx ratio

Write the 5xx ratio query to /root/obs/03-errrate.promql

/root/obs/03-errrate.promql. It is a ratio, so it is a division, and you must apply rate to both the numerator and the denominator for the units to match. If the result is around 0.004 it is normal time, and if it exceeds 0.1 you are looking at an incident period — this environment contains two periods in which errors spike.

p95 latency

Write the p95 latency query to /root/obs/04-p95.promql

/root/obs/04-p95.promql. You cannot compute quantiles from _sum and _count. Apply rate to the buckets, aggregate them while keeping the le label, and then feed that into histogram_quantile. If you leave out by (le), a wrong number comes out without an error — which is why it is more dangerous.

Find the metrics with the most time series

Write the five metrics with the most time series to /root/obs/05-top.txt as 이름 시계열수 (name and series count)

You only know cardinality by measuring it. Throw promq 'topk(5, count by (__name__)({__name__=~".+"}))', and copy the resulting table into /root/obs/05-top.txt as five lines — on each line 이름 시계열수 (name and series count), the largest first. The ranking changes as more targets come up. That is why you leave the table measured at that time, not a single name.

Write the recording rules file

Write the recording rules to /etc/prometheus/rules/labhub.yml

Create /etc/prometheus/rules/labhub.yml. The structure is groups → rules → record/expr. The record name follows the 수준:메트릭:연산 convention (level, metric, operation), as in job:http_requests:rate5m. After saving, be sure to check it with promtool check rules /etc/prometheus/rules/labhub.yml.

Apply the rules and confirm the new time series

Apply the configuration and confirm that the new time series appears

Apply it without restarting Prometheus — curl -X POST http://127.0.0.1:9090/-/reload. Wait about 15 seconds and then check with promq 'job:http_requests:rate5m' whether the new time series has appeared. Rules are computed on every evaluation interval, so it does not appear immediately.

Calculate when the disk will fill up

Write the disk exhaustion forecast query to /root/obs/08-predict.promql

/root/obs/08-predict.promql. predict_linear gives you 'the value N seconds from now if the current trend continues'. Write a query that asks whether node_filesystem_avail_bytes drops below 0 in 6 hours. The disk in this environment has actually been made to decrease monotonically.