TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

The Panel Type Is the Shape of the Question

Continue in TT Lab

In one line

The criterion for choosing a type is not taste. It splits by whether the question is about time, about a single value right now, or about spread.

Why this was needed

If you choose the wrong type, the graph is drawn but the answer does not come out. And that fact is not noticeable — the screen looks fine and the numbers are right, it just cannot show you what you need at the moment you need it. A wrong query shows itself as a blank screen, but a wrong type is silent.

Shape of the question Type Example
How did it change over time timeseries Error rate, p95 latency
A single value right now stat Current request rate, remaining error budget
How full against a fixed upper limit gauge Disk utilization, connection pool
How the values are spread heatmap Distribution of response time
Ranking of several targets table · bar Error count per handler

The moment each one goes wrong

A gauge on requests per second. A gauge is a tool that shows "where between 0 and the maximum." Requests per second has no maximum. Wherever the needle sits, that position tells you nothing, and since there is no time axis, it cannot answer "since when did it grow" either. Use a gauge only for values with a defined upper limit.

One line graph for a distribution. A single average response time line perfectly hides "half of the users waited 2 seconds." An average is dragged around by a few large values yet does not show that those values existed. To see spread, overlay several quantiles or use a heatmap.

Only stat on a health dashboard. From the single number "error rate is 3% now," you cannot tell whether it is rising or falling. In an incident, that difference is everything. When you show a single number, put a panel with a time axis next to it.

A quantile is not an average

When you compute a quantile from a histogram, the le label is the bucket boundary. So you must keep le and sum the rest.

histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

If you leave out sum by (le), the per-instance buckets go their own ways and the values get tangled. Conversely, averaging p95s you have already computed is also wrong — quantiles cannot be averaged. The average of the p95s of ten instances is not the overall p95. What you must combine is not the results but the buckets.

Common misconceptions

"I can change the type later." You do not. A panel, once drawn, stays as it is, and a wrong type quietly makes that panel useless.

"Pie charts look nice." A pie chart has no time, and it is hard to compare slice sizes by eye. For the same data, a table or bars are almost always better.

What really matters in practice

Write the unit and the period in the panel title. Write "5xx ratio (5-minute rate)," not "Errors." The person who got paged is seeing that panel for the first time, and has no way to know other than the title whether the number 0.004 is a ratio, a count, per second, or per minute.

And tell Grafana the unit of the axis (things like percentunit and seconds). 0.4% reads much faster at 3 a.m. than 0.004, and 223ms than 0.223.