Throwing Queries at a Real Prometheus
Goal
You throw queries at a real Prometheus server and read the numbers that come back. By the end of this lab, you will have written rate, ratio, quantile, and forecast queries yourself and will know why they were wrong when they are wrong.
Why it matters
You can memorize PromQL syntax in a day. The hard part is judging whether the number that
came back is right. A histogram_quantile that omits by (le) does not raise an error —
it just gives you a wrong number. If you put that on a dashboard, nobody knows for months.
That is why this environment comes preloaded with 12 hours of time series. There are two periods in which errors spike, one period in which only p99 spikes, and a disk that decreases monotonically. You can check the answer by whether the queries you wrote find those events.
Environment
promq "<PromQL>" 쿼리를 던진다. promq -r 로 원본 JSON
lab-status 무엇이 떠 있는지, 대상이 몇 개인지
tail -40 /var/log/lab-init.log 준비 과정 로그
Prometheus is at 127.0.0.1:9090, the sample service exporter is at 9101, and
node_exporter is at 9100. Start Grafana with lab-start-grafana and view it through
web preview port 3000.
Steps
- Find which labels
http_requests_totalhas and write them to/root/obs/01-labels.txt - Write the per-second request rate query to
/root/obs/02-rate.promql - Write the 5xx ratio query to
/root/obs/03-errrate.promql - Write the p95 latency query to
/root/obs/04-p95.promql - Write the five metrics with the most time series to
/root/obs/05-top.txtas이름 시계열수(name and series count) - Write the recording rules to
/etc/prometheus/rules/labhub.yml - Apply the configuration and confirm that the new time series appears
- Write the disk exhaustion forecast query to
/root/obs/08-predict.promql
Notes
- Before writing a query to a file, throw it with
promqfirst. If the result is empty, a label name is wrong. - If
promq 'up'gives 3, scraping is healthy. - If the result is
NaN, you are looking at a period where the denominator is 0.
First look at what time series exist
Find which labels http_requests_total has and write them to /root/obs/01-labels.txt
Before writing a query, look at what exists. Scan the metric names with promq 'count by (__name__)({__name__=~".+"})' and look at the labels with promq 'http_requests_total'. Write the label names of http_requests_total to /root/obs/01-labels.txt, one per line (names only, not values).
Per-second request rate
Write the per-second request rate query to /root/obs/02-rate.promql
Write the query to /root/obs/02-rate.promql and check it directly with promq "$(cat /root/obs/02-rate.promql)". The scrape interval is 15 seconds, so the range must be at least four times that so there is no moment when only one sample is captured. rate must be inside sum — if you put it outside, counter resets are handled wrongly.
5xx ratio
Write the 5xx ratio query to /root/obs/03-errrate.promql
/root/obs/03-errrate.promql. It is a ratio, so it is a division, and you must apply rate to both the numerator and the denominator for the units to match. If the result is around 0.004 it is normal time, and if it exceeds 0.1 you are looking at an incident period — this environment contains two periods in which errors spike.
p95 latency
Write the p95 latency query to /root/obs/04-p95.promql
/root/obs/04-p95.promql. You cannot compute quantiles from _sum and _count. Apply rate to the buckets, aggregate them while keeping the le label, and then feed that into histogram_quantile. If you leave out by (le), a wrong number comes out without an error — which is why it is more dangerous.
Find the metrics with the most time series
Write the five metrics with the most time series to /root/obs/05-top.txt as 이름 시계열수 (name and series count)
You only know cardinality by measuring it. Throw promq 'topk(5, count by (__name__)({__name__=~".+"}))', and copy the resulting table into /root/obs/05-top.txt as five lines — on each line 이름 시계열수 (name and series count), the largest first. The ranking changes as more targets come up. That is why you leave the table measured at that time, not a single name.
Write the recording rules file
Write the recording rules to /etc/prometheus/rules/labhub.yml
Create /etc/prometheus/rules/labhub.yml. The structure is groups → rules → record/expr. The record name follows the 수준:메트릭:연산 convention (level, metric, operation), as in job:http_requests:rate5m. After saving, be sure to check it with promtool check rules /etc/prometheus/rules/labhub.yml.
Apply the rules and confirm the new time series
Apply the configuration and confirm that the new time series appears
Apply it without restarting Prometheus — curl -X POST http://127.0.0.1:9090/-/reload. Wait about 15 seconds and then check with promq 'job:http_requests:rate5m' whether the new time series has appeared. Rules are computed on every evaluation interval, so it does not appear immediately.
Calculate when the disk will fill up
Write the disk exhaustion forecast query to /root/obs/08-predict.promql
/root/obs/08-predict.promql. predict_linear gives you 'the value N seconds from now if the current trend continues'. Write a query that asks whether node_filesystem_avail_bytes drops below 0 in 6 hours. The disk in this environment has actually been made to decrease monotonically.