All six panels were answering different questions
Goal
You receive a dashboard with defects, fix the calculation, null handling, stacking, number of series, and title and description, and build a checker that catches the same defects in the next dashboard too and make it pass.
Why it matters
The place where a dashboard is quietly wrong is not the query but the panel options. If the query is right but a one-number panel is showing an average, a 20-minute spike disappears in the six-hour average. If you connect missing points, the very fact that collection was cut is erased, and if you stack series, the top line gets read as an individual value. All three are not "wrong values" but "accurate answers to questions nobody asked," so the person looking cannot tell they are misreading. So you pin down in the description which question each panel answers, and make a checker enforce that promise.
Steps
- Start Grafana with
lab-start-grafanaand upload the defective dashboard/opt/lab/gfd/gfd-misread/broken.jsonto Grafana as it is, without editing it (the uid isgfd-misreadas written in the file, with 6 panels). You can POST it to/api/dashboards/dbwithcurl. - Panel 1 (
요청률, the Korean title means "request rate") currently shows an average. Pick a past range (at least 1 hour, ending before now), measure directly the last value, average, and maximum ofsum(rate(http_requests_total{job="shop-api"}[5m]))over that range, and write them to/root/gfd-misread/02-calc.txtas five linesstart=end=last=mean=max=(start and end are epoch seconds). Then change panel 1's calculation tolastNotNulland save. - Write three lines
connected=none=zero=to/root/gfd-misread/03-null.txt. Each line must explain in at least 40 characters what that null handling choice claims to the person looking, and the three lines must differ from each other. Then changespanNullsof panel 2 (대기열, the Korean title means "queue") tofalseand save. - Panel 3 (
핸들러별 요청률, the Korean title means "request rate by handler") draws the series stacked. Pick one past time, measure the overall total at that moment and the value of the single/api/ordershandler, and write them to/root/gfd-misread/04-stack.txtas three linesat=total=orders=(at is epoch seconds). Then change panel 3'sstacking.modetononeand save. - The query of panel 4 (
핸들러 지연, the Korean title means "handler latency") returns one series per handler. Count that number and write it tobefore=of/root/gfd-misread/05-series.txt, fix the query so that this panel answers the single question "what is the p95 of the slowest handler right now" and save it, and then write the number of series of the fixed query toafter=. - Fix the titles of all six panels so that you can tell what the panel looks at (names like
그래프, the Korean word for "graph," are forbidden), and in the descriptions write the question sentence that panel answers, ending in a question mark (at least 12 characters, different for each panel). Save the fixed dashboard. - Create
/root/gfd-misread/lint.py. It takes the dashboard JSON file path as an argument and prints the violations of the five rules below, one per line (starting withR1toR5), and must end with exit code 1 if there is even one violation. R1 the calculation of a stat panel contains the average · R2spanNullsis true · R3stacking.modeisnormal· R4 the description does not end in a question mark · R5 the title is empty or is그래프,패널, or차트(the Korean words for "graph," "panel," and "chart"). Run it on the original/opt/lab/gfd/gfd-misread/broken.jsonand confirm that all five rules are caught. - Download the dashboard now in Grafana as it stands and save it to
/root/gfd-misread/fixed.json(only the.dashboardbody), and run the checker to confirm 0 violations and exit code 0. Then write to/root/gfd-misread/08-review.md, as five linesR1=throughR5=, what you fixed and why, each at least 30 characters.
Notes
- Start Grafana with
lab-start-grafana(it takes a few seconds). Also check it by eye on port 3000 of the web preview. - The defective original dashboard is in
/opt/lab/gfd/gfd-misread/broken.json. Do not edit this file; only read it — you use it again in step 8. - The API for saving a dashboard is
POST /api/dashboards/dband the body is{"dashboard": ..., "overwrite": true}. - For a range query use
/api/v1/query_range, and for a single-moment value sendtime=along with/api/v1/query. Both paths can be attached after/api/datasources/proxy/uid/<uid>/to throw through Grafana. - Common mistake 1: not saving after fixing. Even if you change it on screen, if you do not save it through the API, what the grader sees is the old version.
- Common mistake 2: basing the range on
지금(the Korean word for "now"). Only if you pin it to a past absolute time do you get the same value when you measure again. - Panel options · Standard options · Dashboard JSON model · Dashboard HTTP API · Prometheus query API
Put the state before fixing on screen
Start Grafana with lab-start-grafana and upload the defective dashboard /opt/lab/gfd/gfd-misread/broken.json to Grafana as it is, without editing it (the uid is gfd-misread as written in the file, with 6 panels). You can POST it to /api/dashboards/db with curl.
The save API takes the dashboard body inside the dashboard key and sends overwrite along with it. jq -n --slurpfile is convenient for building that shape from the file. Grafana takes a few seconds to come up, so first check whether /api/health responds.
What is a one-number panel calculating
Panel 1 (요청률, the Korean title means "request rate") currently shows an average. Pick a past range (at least 1 hour, ending before now), measure directly the last value, average, and maximum of sum(rate(http_requests_total{job="shop-api"}[5m])) over that range, and write them to /root/gfd-misread/02-calc.txt as five lines start= end= last= mean= max= (start and end are epoch seconds). Then change panel 1's calculation to lastNotNull and save.
A range query is /api/v1/query_range and you send start, end, and step along with it. If you throw it through the data source proxy, it goes by the same path Grafana uses. The calculation is in options.reduceOptions.calcs of the dashboard JSON. The reason to measure over a fixed range is so that you get the same value when you measure again later.
Connecting missing points is a claim too
Write three lines connected= none= zero= to /root/gfd-misread/03-null.txt. Each line must explain in at least 40 characters what that null handling choice claims to the person looking, and the three lines must differ from each other. Then change spanNulls of panel 2 (대기열, the Korean title means "queue") to false and save.
If you connect the line across a stretch where collection was cut, it looks as if there was a value during that time too. Null handling is in fieldConfig.defaults.custom.spanNulls. Also think about in what situation each of the three choices is right — it is not a problem with a single right answer.
The top line of a stack is not any series
Panel 3 (핸들러별 요청률, the Korean title means "request rate by handler") draws the series stacked. Pick one past time, measure the overall total at that moment and the value of the single /api/orders handler, and write them to /root/gfd-misread/04-stack.txt as three lines at= total= orders= (at is epoch seconds). Then change panel 3's stacking.mode to none and save.
For a single-moment value, send time= along with /api/v1/query. You must pin down the time so that you get the same value when you measure again later. The stacking setting is in fieldConfig.defaults.custom.stacking.mode. If you compare the two numbers, you can see how wrong you are when you read the top line as an individual series.
When four series come to a panel that asks for one value
The query of panel 4 (핸들러 지연, the Korean title means "handler latency") returns one series per handler. Count that number and write it to before= of /root/gfd-misread/05-series.txt, fix the query so that this panel answers the single question "what is the p95 of the slowest handler right now" and save it, and then write the number of series of the fixed query to after=.
If you throw it with promq "<쿼리>" (the placeholder is the query), the result series come out one per line. There is an aggregation operator that keeps only the largest value among several series. The query is targets[0].expr in the dashboard JSON.
The title and description are that panel's question
Fix the titles of all six panels so that you can tell what the panel looks at (names like 그래프, the Korean word for "graph," are forbidden), and in the descriptions write the question sentence that panel answers, ending in a question mark (at least 12 characters, different for each panel). Save the fixed dashboard.
The description is the description that each panel in the dashboard JSON has. On screen it appears as the info mark next to the panel title. If you write it as a question sentence, half a year later you can judge whether it is safe to delete this panel — you only have to see whether that question is still being asked.
Catch the same defects in the next dashboard too
Create /root/gfd-misread/lint.py. It takes the dashboard JSON file path as an argument and prints the violations of the five rules below, one per line (starting with R1 to R5), and must end with exit code 1 if there is even one violation. R1 the calculation of a stat panel contains the average · R2 spanNulls is true · R3 stacking.mode is normal · R4 the description does not end in a question mark · R5 the title is empty or is 그래프, 패널, or 차트 (the Korean words for "graph," "panel," and "chart"). Run it on the original /opt/lab/gfd/gfd-misread/broken.json and confirm that all five rules are caught.
If you leave the rules written as prose, the next dashboard is born in the same state again. Write them as running code. Do not forget that a panel folded inside a row is also a panel. The file may be wrapped as {"dashboard": ...} or may be the dashboard body itself.
The fixed dashboard passes its own checker
Download the dashboard now in Grafana as it stands and save it to /root/gfd-misread/fixed.json (only the .dashboard body), and run the checker to confirm 0 violations and exit code 0. Then write to /root/gfd-misread/08-review.md, as five lines R1= through R5=, what you fixed and why, each at least 30 characters.
To keep the file and the screen from drifting apart, you must download again after fixing. If the checker does not end with 0, the output tells you which panel remains. The record is left so that the next person does not have to make the same judgment again.