TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

The panel was accurate; the question was different

Continue in TT Lab

In one line

It is rare for a panel to draw a wrong value. What is common is a panel accurately showing the answer to a question we did not ask.

Why this was needed

There is a sentence that often comes up in a retrospective after an outage. "But the dashboard looked normal." Yet if you measure the raw data of the same period again, there was clearly an anomaly. The query was right and the data was there. What got lost in between is the panel options.

Suppose there is a panel that shows one number large. People read it as "the current value." But what that panel actually computes may be the average of the time range visible on screen. If you were viewing six hours, a 20-minute spike almost disappears in the average. The panel was accurate. It was just answering a question nobody asked: "what is the average over the past six hours."

The same goes for missing points. If collection is cut for 30 minutes, there are no points in that stretch. If you connect the line as is, the graph flows smoothly, and the person looking reads that the service was running fine during those 30 minutes too. What was cut was the fact that there is no data, and the choice to connect the line erases that fact.

How it works

A Grafana panel has roughly three layers. The query fetches the time series, field settings and transformations refine it, and the visualization options decide how to draw it. Panels that get misread mostly come from the second and third layers. That is why the cause is invisible no matter how hard you look at the query.

Option What it decides If set wrongly
Calculation (reduceOptions.calcs) How to reduce one time series to one number With average, spikes get buried
Null handling (spanNulls) Whether to connect missing points If connected, collection gaps are erased
Stacking (stacking.mode) Whether to stack the series If stacked, individual values cannot be read
Number of series How many one panel shows A question about one value gets four boxes

Start with the calculation. A time series has many points, but a one-number panel has to pick one of them. If you pick the last value, it is "now," if you pick the average, it is "the average over this period," and if you pick the maximum, it is "the worst of this period." The three are completely different questions, and the panel title does not distinguish among them. So if the title is "request rate," people will surely read it as "the current request rate." If that reading and the panel's calculation are off, the screen is quietly wrong.

Null handling has three branches: connect, break, or fill with 0. The three each claim "there was a value in between," "what happened in between is unknown," and "it was 0 in between." A value that truly cannot be known when collection is cut, like the rate of change of a counter, should be drawn broken, and for a value where 0 is meaningful, like queue length, it depends on the situation. What matters is to choose while knowing that either way you are making a claim.

Stacking is right only when you want to see the overall total. The top line of a stacked graph is not the value of any series but the sum of everything, yet the human eye reads the top line as "the largest series." If the goal is to compare series with each other, you must not stack, and if you are curious about proportions, use a shape that states the ratio explicitly, like a 100% stack.

Finally, the title and description. A Grafana panel has a place to write a description (panel editing documentation), and if you write a question sentence there, two things are solved at once. The reader learns how to read this panel, and half a year later you can also judge whether it is safe to delete this panel — you only have to see whether that question is still being asked.

What it looks like in the field

This actually happened on a payment service's health dashboard. The error rate panel was green all day, yet customer inquiries kept coming in. The panel was a stat, its calculation was the average, and its default time range was 24 hours. In the morning the error rate was 30% for 12 minutes, but as a daily average it was 0.25%. When the calculation was changed to the last value and the time range was cut to one hour, the same incident caught the eye within 3 minutes the following week.

Another was stacking. The request-rate-by-handler panel was drawn stacked, and the team read the top line as the request rate of /api/orders and made a capacity plan. The real /api/orders was half of that. The misunderstanding went away not by fixing the numbers but just by turning off the stacking.

What can and cannot be judged in this environment

The Grafana in this Pod really runs, but it has no image renderer plugin. So how a panel is drawn on screen cannot be inspected. Instead, judgment is made by the dashboard JSON model (calculation, null handling, stacking, title, description) and by the values from actually running its queries against the data source. How colors actually look and where lines break you ultimately have to see by eye once — in the lab too, we recommend opening port 3000 in the web preview to check.

What you will do in the next lab

You upload a dashboard containing six defects as it is, and fix the calculation, null handling, stacking, number of series, and title and description one by one. Each time you fix something, you confirm why that choice was wrong with numbers — measuring last, mean, and max directly over a fixed range, and measuring the sum and an individual value at one time to see what the stacking hid. At the end, you build a checker that catches the same defects in the next dashboard too, and confirm that the fixed dashboard passes that checker.