The same outage was invisible on the dashboard
Goal
You ask the same question of the same 12 hours of data, changing only the rate window and the step, measure how the answer differs in numbers, and fix and submit three broken panels copied from a production dashboard.
Why it matters
Trusting a dashboard should mean knowing what summary each panel is making. But for most panels, once someone has built them, nobody ever opens the query again. If the window is one hour, a 20-minute incident shows up squashed to a third of its height, and if the window is shorter than the scrape interval, the screen shows 'no data' even though the data exists. If you average ratios, one hour at 1 a.m. and one hour at midday get the same weight, and if you average quantiles, you get a number that means nothing. The purpose of this lab is not to learn more PromQL but to become able to look at a single panel and ask 'what did this picture erase?'
Steps
- Create
/root/obs-graph-traps/qr.py. When called aspython3 qr.py '<PromQL>' <step초> [<범위 시간>] [<임계값>](the query, the step in seconds, the range in hours, and the threshold), it throws one request to/api/v1/query_rangewithstart=지금-범위,end=지금, andstep=step초(start is now minus the range, end is now, and step is the step in seconds) and prints one line,points=<돌아온 값의 총 개수> max=<최댓값> over=<임계값보다 큰 값의 개수>(the total number of values returned, the maximum, and the number of values greater than the threshold). The default range is 12 (hours) and the default threshold is 0.01. If there are no values at all, printpoints=0 max=NA over=0. Write the maximum to six decimal places. - Throw the expression that asks for the 5xx ratio,
sum(rate(http_requests_total{job="shop-api",status="500"}[창])) / sum(rate(http_requests_total{job="shop-api"}[창]))(where the placeholder is the window), four times, changing only the window. The windows are15s,30s,1m, and5m, the step is 60, and the range is 12 hours. In/root/obs-graph-traps/window-points.tsv, write four lines with no header, and each line has two tab-separated columns,<창> <점 수>(the window and the number of points). Then write one line starting withreason=of at least 40 characters in/root/obs-graph-traps/02-why.txtexplaining why there are no points in the shortest window. - Throw the same 5xx ratio expression five times with the windows
1m,5m,15m,30m, and1h(step 60, range 12 hours, threshold 0.01). In/root/obs-graph-traps/window-shape.tsv, write five lines with no header, and each line has three tab-separated columns,<창> <최댓값> <0.01 을 넘은 점 수>(the window, the maximum, and the number of points exceeding 0.01). The maximum is to six decimal places and the number of points is an integer. - Grafana calculates
$__rate_intervalasmax($__interval + 스크레이프간격, 4 * 스크레이프간격)(where the placeholder is the scrape interval). This Pod's scrape interval is 15 seconds. Calculate the rate window for steps of15,108,600, and3600seconds, and then throw the 5xx ratio expression with that window at the same step (range 12 hours). In/root/obs-graph-traps/rate-interval.tsv, write four lines with no header, and each line has four tab-separated columns,<step초> <rate 창 초> <점 수> <최댓값>(the step in seconds, the rate window in seconds, the number of points, and the maximum). Then write one line starting withreason=of at least 40 characters in/root/obs-graph-traps/04-note.txtexplaining why, only at step 3600, the maximum differs each time you throw the same query again. - Get the 5xx ratio of the same 12 hours in two ways. One is the simple average of the 1-minute-interval ratios,
avg_over_time((sum(rate(http_requests_total{job="shop-api",status="500"}[5m])) / sum(rate(http_requests_total{job="shop-api"}[5m])))[12h:1m]), and the other is the traffic-weighted overall ratio,sum(increase(http_requests_total{job="shop-api",status="500"}[12h])) / sum(increase(http_requests_total{job="shop-api"}[12h])). Write three lines in/root/obs-graph-traps/weighted.txt—avg_of_ratio=<소수 여섯 자리>,weighted_ratio=<소수 여섯 자리>, andgap_pct=<소수 두 자리>(six decimal places, six decimal places, and two decimal places).gap_pctis (the first value − the second value) ÷ the second value × 100. - Throw
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[창])))(where the placeholder is the window) three times with the windows5m,15m, and12h(step 60, range 12 hours). In/root/obs-graph-traps/p99-window.tsv, write three lines with no header, and each line has two tab-separated columns,<창> <12시간 최댓값>(the window and the 12-hour maximum). Then write one line starting withreason=of at least 40 characters in/root/obs-graph-traps/06-note.txtexplaining why a long-window quantile comes out smaller than the maximum of the short-window quantiles. - Panel
p1in/opt/lab/graph/panels.ymlasks for 'the 5xx count over the last 12 hours' but produces a value with a decimal point. In/root/obs-graph-traps/fix-count.promql, write a PromQL that gives the 5xx count for the same 12 hours as an integer (you may leave comments with#). Then write two lines in/root/obs-graph-traps/count-panel.txt—broken_value=<소수 세 자리>(three decimal places) is the value that the p1 expression gives now, andfixed_value=<정수>(an integer) is the value the fixed query gives. The grader actually throws the fixed query and looks at the value. - Panels
p2(error ratio) andp3(p99 latency) in/opt/lab/graph/panels.ymlcannot answer 'how high was it at its highest'. Write the fixed queries to/root/obs-graph-traps/fix-ratio.promqland/root/obs-graph-traps/fix-p99.promql. Then write two lines with no header in/root/obs-graph-traps/verdict.tsv, and each line has four tab-separated columns,<패널id> <고치기 전 12시간 최댓값> <고친 뒤 12시간 최댓값> <무엇이 틀렸었나>(the panel id, the 12-hour maximum before the fix, the 12-hour maximum after the fix, and what was wrong). The first line isp2and the second isp3, the maximums are measured with step 60 and a 12-hour range to six decimal places. The fourth column must be at least 20 characters and include at least one number.
Notes
- The working directory is
/root/obs-graph-traps. If it does not exist, create it first. - The original panels are in
/opt/lab/graph/panels.yml. You use theexprof this file in steps 7 and 8. - You can throw a query directly with
promq "<PromQL>"(the placeholder is the query). For range queries, use theqr.pyyou built in step 1. - This Pod's scrape interval is 15 seconds, the resolution of the backfilled period is 30 seconds, and the data covers 12 hours.
- Do not start Grafana. This lab looks only at the queries behind the panels — building dashboards is covered in another lab.
- Common mistake: choosing the window by looking only at the maximum. If you shorten the window the peak survives, but the noise comes back with it.
- Common mistake: reading the result of
increase()as an exact count. It is an estimate extrapolated at both ends. - HTTP API — range queries · Query functions · Histograms and summaries · Grafana — $__rate_interval · Monitoring Distributed Systems (SRE Book, Chapter 6)
Build a tool that summarizes a range query as numbers
Create /root/obs-graph-traps/qr.py. When called as python3 qr.py '<PromQL>' <step초> [<범위 시간>] [<임계값>] (the query, the step in seconds, the range in hours, and the threshold), it throws one request to /api/v1/query_range with start=지금-범위, end=지금, and step=step초 (start is now minus the range, end is now, and step is the step in seconds) and prints one line, points=<돌아온 값의 총 개수> max=<최댓값> over=<임계값보다 큰 값의 개수> (the total number of values returned, the maximum, and the number of values greater than the threshold). The default range is 12 (hours) and the default threshold is 0.01. If there are no values at all, print points=0 max=NA over=0. Write the maximum to six decimal places.
The range response has the values as strings in data.result[].values[][1]. If there are several time series, count them all together. A NaN string may be mixed in, so you must filter it with math.isnan so the maximum is not ruined. Use only the standard library (urllib.request, json, time, math).
An empty graph does not mean there is no traffic
Throw the expression that asks for the 5xx ratio, sum(rate(http_requests_total{job="shop-api",status="500"}[창])) / sum(rate(http_requests_total{job="shop-api"}[창])) (where the placeholder is the window), four times, changing only the window. The windows are 15s, 30s, 1m, and 5m, the step is 60, and the range is 12 hours. In /root/obs-graph-traps/window-points.tsv, write four lines with no header, and each line has two tab-separated columns, <창> <점 수> (the window and the number of points). Then write one line starting with reason= of at least 40 characters in /root/obs-graph-traps/02-why.txt explaining why there are no points in the shortest window.
This Pod's scrape interval is 15 seconds. rate can compute a rate of change only if there are two or more samples in the window, and if there is only one, it returns not an error but an empty result. That is why the official documentation says to set the window to at least four times the scrape interval. The resolution of the backfilled period is 30 seconds, so even a 30-second window is almost empty.
As you widen the window, the peak gets lower and the width gets wider
Throw the same 5xx ratio expression five times with the windows 1m, 5m, 15m, 30m, and 1h (step 60, range 12 hours, threshold 0.01). In /root/obs-graph-traps/window-shape.tsv, write five lines with no header, and each line has three tab-separated columns, <창> <최댓값> <0.01 을 넘은 점 수> (the window, the maximum, and the number of points exceeding 0.01). The maximum is to six decimal places and the number of points is an integer.
Since you fixed step at 60, one point is exactly 1 minute. You can read the third column as 'the incident length the graph tells you (in minutes)'. Look at both directions together: how many times smaller the height gets while the width grows by how many times.
The panel width decides the window — calculate $__rate_interval by hand
Grafana calculates $__rate_interval as max($__interval + 스크레이프간격, 4 * 스크레이프간격) (where the placeholder is the scrape interval). This Pod's scrape interval is 15 seconds. Calculate the rate window for steps of 15, 108, 600, and 3600 seconds, and then throw the 5xx ratio expression with that window at the same step (range 12 hours). In /root/obs-graph-traps/rate-interval.tsv, write four lines with no header, and each line has four tab-separated columns, <step초> <rate 창 초> <점 수> <최댓값> (the step in seconds, the rate window in seconds, the number of points, and the maximum). Then write one line starting with reason= of at least 40 characters in /root/obs-graph-traps/04-note.txt explaining why, only at step 3600, the maximum differs each time you throw the same query again.
$__interval is the panel's time range divided by its pixel width, so here the step is $__interval. 108 is the value you get when drawing a 12-hour panel at 400 pixels, and 3600 comes up when you look at a much wider range. You can put the window in as a string in seconds as it is, like 3615s. The grid starts at start and is placed at step intervals, and start moves along with 'now' — when the step gets longer than the incident, where the grid points fall on the peak changes the maximum.
The average of ratios is not the overall ratio
Get the 5xx ratio of the same 12 hours in two ways. One is the simple average of the 1-minute-interval ratios, avg_over_time((sum(rate(http_requests_total{job="shop-api",status="500"}[5m])) / sum(rate(http_requests_total{job="shop-api"}[5m])))[12h:1m]), and the other is the traffic-weighted overall ratio, sum(increase(http_requests_total{job="shop-api",status="500"}[12h])) / sum(increase(http_requests_total{job="shop-api"}[12h])). Write three lines in /root/obs-graph-traps/weighted.txt — avg_of_ratio=<소수 여섯 자리>, weighted_ratio=<소수 여섯 자리>, and gap_pct=<소수 두 자리> (six decimal places, six decimal places, and two decimal places). gap_pct is (the first value − the second value) ÷ the second value × 100.
You can throw both values directly with promq. The first expression counts a quiet 1 minute at dawn and a busy 1 minute at midday with the same weight, and the second gives weight in proportion to the number of requests. If the incident happened during a quiet time, think first about which side will be larger.
A long-window p99 is not the worst p99
Throw histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[창]))) (where the placeholder is the window) three times with the windows 5m, 15m, and 12h (step 60, range 12 hours). In /root/obs-graph-traps/p99-window.tsv, write three lines with no header, and each line has two tab-separated columns, <창> <12시간 최댓값> (the window and the 12-hour maximum). Then write one line starting with reason= of at least 40 characters in /root/obs-graph-traps/06-note.txt explaining why a long-window quantile comes out smaller than the maximum of the short-window quantiles.
Histogram buckets are also counters, so the rate window applies as is. If the window is long, slow requests and fast requests get mixed in one pot and the tail is diluted. That is why a single 'p99 over the last 12 hours' panel cannot answer 'how high was it at its worst'. You can leave the threshold at its default — all you look at in this step is the maximum.
Application 1 — the count panel shows a decimal point
Panel p1 in /opt/lab/graph/panels.yml asks for 'the 5xx count over the last 12 hours' but produces a value with a decimal point. In /root/obs-graph-traps/fix-count.promql, write a PromQL that gives the 5xx count for the same 12 hours as an integer (you may leave comments with #). Then write two lines in /root/obs-graph-traps/count-panel.txt — broken_value=<소수 세 자리> (three decimal places) is the value that the p1 expression gives now, and fixed_value=<정수> (an integer) is the value the fixed query gives. The grader actually throws the fixed query and looks at the value.
increase() and rate() extrapolate at both ends because samples do not sit exactly on the window boundary. That is why a count that should be an integer comes out as a decimal. If you want the exact count, you must ask in a way that does not extrapolate — think of the difference between the counter's current value and its value 12 hours ago. The offset modifier helps.
Application 2 — fix and submit the error ratio and p99 panels
Panels p2 (error ratio) and p3 (p99 latency) in /opt/lab/graph/panels.yml cannot answer 'how high was it at its highest'. Write the fixed queries to /root/obs-graph-traps/fix-ratio.promql and /root/obs-graph-traps/fix-p99.promql. Then write two lines with no header in /root/obs-graph-traps/verdict.tsv, and each line has four tab-separated columns, <패널id> <고치기 전 12시간 최댓값> <고친 뒤 12시간 최댓값> <무엇이 틀렸었나> (the panel id, the 12-hour maximum before the fix, the 12-hour maximum after the fix, and what was wrong). The first line is p2 and the second is p3, the maximums are measured with step 60 and a 12-hour range to six decimal places. The fourth column must be at least 20 characters and include at least one number.
p2's expression is not wrong; its window is too long. p3 is wrong in two places — the window is long, and it averages quantiles that have already been computed. For a histogram you must first aggregate the buckets by le and then compute the quantile. It is normal for the fixed values to be the same as the numbers you already saw in steps 3 and 6.