TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

All six panels were answering different questions

Continue in TT Lab

Goal

You receive a dashboard with defects, fix the calculation, null handling, stacking, number of series, and title and description, and build a checker that catches the same defects in the next dashboard too and make it pass.

Why it matters

The place where a dashboard is quietly wrong is not the query but the panel options. If the query is right but a one-number panel is showing an average, a 20-minute spike disappears in the six-hour average. If you connect missing points, the very fact that collection was cut is erased, and if you stack series, the top line gets read as an individual value. All three are not "wrong values" but "accurate answers to questions nobody asked," so the person looking cannot tell they are misreading. So you pin down in the description which question each panel answers, and make a checker enforce that promise.

Steps

  1. Start Grafana with lab-start-grafana and upload the defective dashboard /opt/lab/gfd/gfd-misread/broken.json to Grafana as it is, without editing it (the uid is gfd-misread as written in the file, with 6 panels). You can POST it to /api/dashboards/db with curl.
  2. Panel 1 (요청률, the Korean title means "request rate") currently shows an average. Pick a past range (at least 1 hour, ending before now), measure directly the last value, average, and maximum of sum(rate(http_requests_total{job="shop-api"}[5m])) over that range, and write them to /root/gfd-misread/02-calc.txt as five lines start= end= last= mean= max= (start and end are epoch seconds). Then change panel 1's calculation to lastNotNull and save.
  3. Write three lines connected= none= zero= to /root/gfd-misread/03-null.txt. Each line must explain in at least 40 characters what that null handling choice claims to the person looking, and the three lines must differ from each other. Then change spanNulls of panel 2 (대기열, the Korean title means "queue") to false and save.
  4. Panel 3 (핸들러별 요청률, the Korean title means "request rate by handler") draws the series stacked. Pick one past time, measure the overall total at that moment and the value of the single /api/orders handler, and write them to /root/gfd-misread/04-stack.txt as three lines at= total= orders= (at is epoch seconds). Then change panel 3's stacking.mode to none and save.
  5. The query of panel 4 (핸들러 지연, the Korean title means "handler latency") returns one series per handler. Count that number and write it to before= of /root/gfd-misread/05-series.txt, fix the query so that this panel answers the single question "what is the p95 of the slowest handler right now" and save it, and then write the number of series of the fixed query to after=.
  6. Fix the titles of all six panels so that you can tell what the panel looks at (names like 그래프, the Korean word for "graph," are forbidden), and in the descriptions write the question sentence that panel answers, ending in a question mark (at least 12 characters, different for each panel). Save the fixed dashboard.
  7. Create /root/gfd-misread/lint.py. It takes the dashboard JSON file path as an argument and prints the violations of the five rules below, one per line (starting with R1 to R5), and must end with exit code 1 if there is even one violation. R1 the calculation of a stat panel contains the average · R2 spanNulls is true · R3 stacking.mode is normal · R4 the description does not end in a question mark · R5 the title is empty or is 그래프, 패널, or 차트 (the Korean words for "graph," "panel," and "chart"). Run it on the original /opt/lab/gfd/gfd-misread/broken.json and confirm that all five rules are caught.
  8. Download the dashboard now in Grafana as it stands and save it to /root/gfd-misread/fixed.json (only the .dashboard body), and run the checker to confirm 0 violations and exit code 0. Then write to /root/gfd-misread/08-review.md, as five lines R1= through R5=, what you fixed and why, each at least 30 characters.

Notes

Put the state before fixing on screen

Start Grafana with lab-start-grafana and upload the defective dashboard /opt/lab/gfd/gfd-misread/broken.json to Grafana as it is, without editing it (the uid is gfd-misread as written in the file, with 6 panels). You can POST it to /api/dashboards/db with curl.

The save API takes the dashboard body inside the dashboard key and sends overwrite along with it. jq -n --slurpfile is convenient for building that shape from the file. Grafana takes a few seconds to come up, so first check whether /api/health responds.

What is a one-number panel calculating

Panel 1 (요청률, the Korean title means "request rate") currently shows an average. Pick a past range (at least 1 hour, ending before now), measure directly the last value, average, and maximum of sum(rate(http_requests_total{job="shop-api"}[5m])) over that range, and write them to /root/gfd-misread/02-calc.txt as five lines start= end= last= mean= max= (start and end are epoch seconds). Then change panel 1's calculation to lastNotNull and save.

A range query is /api/v1/query_range and you send start, end, and step along with it. If you throw it through the data source proxy, it goes by the same path Grafana uses. The calculation is in options.reduceOptions.calcs of the dashboard JSON. The reason to measure over a fixed range is so that you get the same value when you measure again later.

Connecting missing points is a claim too

Write three lines connected= none= zero= to /root/gfd-misread/03-null.txt. Each line must explain in at least 40 characters what that null handling choice claims to the person looking, and the three lines must differ from each other. Then change spanNulls of panel 2 (대기열, the Korean title means "queue") to false and save.

If you connect the line across a stretch where collection was cut, it looks as if there was a value during that time too. Null handling is in fieldConfig.defaults.custom.spanNulls. Also think about in what situation each of the three choices is right — it is not a problem with a single right answer.

The top line of a stack is not any series

Panel 3 (핸들러별 요청률, the Korean title means "request rate by handler") draws the series stacked. Pick one past time, measure the overall total at that moment and the value of the single /api/orders handler, and write them to /root/gfd-misread/04-stack.txt as three lines at= total= orders= (at is epoch seconds). Then change panel 3's stacking.mode to none and save.

For a single-moment value, send time= along with /api/v1/query. You must pin down the time so that you get the same value when you measure again later. The stacking setting is in fieldConfig.defaults.custom.stacking.mode. If you compare the two numbers, you can see how wrong you are when you read the top line as an individual series.

When four series come to a panel that asks for one value

The query of panel 4 (핸들러 지연, the Korean title means "handler latency") returns one series per handler. Count that number and write it to before= of /root/gfd-misread/05-series.txt, fix the query so that this panel answers the single question "what is the p95 of the slowest handler right now" and save it, and then write the number of series of the fixed query to after=.

If you throw it with promq "<쿼리>" (the placeholder is the query), the result series come out one per line. There is an aggregation operator that keeps only the largest value among several series. The query is targets[0].expr in the dashboard JSON.

The title and description are that panel's question

Fix the titles of all six panels so that you can tell what the panel looks at (names like 그래프, the Korean word for "graph," are forbidden), and in the descriptions write the question sentence that panel answers, ending in a question mark (at least 12 characters, different for each panel). Save the fixed dashboard.

The description is the description that each panel in the dashboard JSON has. On screen it appears as the info mark next to the panel title. If you write it as a question sentence, half a year later you can judge whether it is safe to delete this panel — you only have to see whether that question is still being asked.

Catch the same defects in the next dashboard too

Create /root/gfd-misread/lint.py. It takes the dashboard JSON file path as an argument and prints the violations of the five rules below, one per line (starting with R1 to R5), and must end with exit code 1 if there is even one violation. R1 the calculation of a stat panel contains the average · R2 spanNulls is true · R3 stacking.mode is normal · R4 the description does not end in a question mark · R5 the title is empty or is 그래프, 패널, or 차트 (the Korean words for "graph," "panel," and "chart"). Run it on the original /opt/lab/gfd/gfd-misread/broken.json and confirm that all five rules are caught.

If you leave the rules written as prose, the next dashboard is born in the same state again. Write them as running code. Do not forget that a panel folded inside a row is also a panel. The file may be wrapped as {"dashboard": ...} or may be the dashboard body itself.

The fixed dashboard passes its own checker

Download the dashboard now in Grafana as it stands and save it to /root/gfd-misread/fixed.json (only the .dashboard body), and run the checker to confirm 0 violations and exit code 0. Then write to /root/gfd-misread/08-review.md, as five lines R1= through R5=, what you fixed and why, each at least 30 characters.

To keep the file and the screen from drifting apart, you must download again after fixing. If the checker does not end with 0, the output tells you which panel remains. The record is left so that the next person does not have to make the same judgment again.