Nobody Looks at a Wall of Graphs
In one line
A dashboard is not a picture but a tool that answers one question. Adding panels and getting answers quickly pull in opposite directions.
Why this was needed
There is one reason to open a dashboard during an outage — to find out within 30 seconds "what is not normal right now." But most dashboards cannot answer that question. There are forty panels, thirty-eight of them look the same as usual, and a person has to find with their eyes which one differs from usual. The person doing that at 3 a.m. ends up closing the dashboard and looking at the logs.
The way panels multiply is also predictable. Someone says "it would be nice to see this too" and adds one. Nobody deletes any. To delete one you have to prove "nobody looks at this," and a dashboard has no way to prove that. So a dashboard grows in only one direction.
If you write the question down first, that flow is broken. If you cannot write, in one sentence, "which question does this panel answer" for each panel, that panel can be deleted. Having grounds for deleting is the real reason to write the questions down.
How it works
Build in the order question → panel. If you do it the other way around, it becomes "we have this metric, so let's draw it," and nobody can interpret panels added that way.
| Question | Panel | Why this metric |
|---|---|---|
| Are we returning failures to users right now | 5xx ratio | A count grows along with traffic. What users experience is a probability, which is a ratio |
| Did it get slower | p95 response time | The average hides the slow half |
| How much is coming in | Requests per second | Even if the error rate drops, it is not normal if traffic is 0 |
| Are resources about to run out | Disk and connection utilization | Only values with an upper limit belong here |
These four are the four golden signals (latency, traffic, errors, saturation) of Google's SRE book. When you build a dashboard for a new service, starting with these four is usually right.
Dashboards are not one kind
There are three kinds that, if mixed, will do none of the three jobs.
| Kind | Question it answers | Who looks | Number of panels |
|---|---|---|---|
| Health | Is it normal right now | The person who got paged | 4–6 |
| Diagnosis | Why is it sick | The person looking for the cause | Can be many |
| Capacity | When do we need to scale | The person planning | Viewed weekly |
The moment you mix diagnosis panels into a health dashboard, that dashboard cannot give an answer within 30 seconds. Move diagnosis panels to another dashboard and connect them with a link.
Common misconceptions
"The more you show, the safer." It is the opposite. The more signals on screen, the higher the chance of missing the one that is off. If a health dashboard has more than six panels, diagnosis panels have usually crept in.
"Put CPU first." CPU at 80% means nothing to users. What users experience (errors, latency) comes first, and resources come after. CPU is material for a diagnosis dashboard that asks "why."
What makes a single panel readable
Even with the same data, depending on how it is drawn, it can be read in 30 seconds or you can stare at it for a long time and still not get it. A panel on a health dashboard needs at least these four.
A baseline is visible alongside. With only the current value, you cannot tell whether it is good or bad. Draw the target as a horizontal line, or overlay the same time last week. For a metric with a clear daily cycle like traffic in particular, overlaying yesterday or last week is far more useful than the absolute value.
Units and ranges are honest. If the axis does not start at 0, a small change looks like a cliff, and conversely an overly wide range flattens real changes. Show ratios as percentages, times in milliseconds or seconds, and sizes in bytes, so that people do not have to do mental arithmetic.
Color means state. If you paint several lines in rainbow colors, it means nothing. Keep normal in calm colors and let only problem states stand out. And use red only when it is really bad. If there is one thing that is red even in normal times, the whole dashboard becomes desensitized to red.
The time range fits the question. If the default range of a screen that asks whether things are normal right now is 30 days, changes in the last few minutes get squashed into a single point. A health dashboard is set short (around an hour), and a capacity dashboard long.
One more thing is recommended. Put a description on the panel. If you write in a line or two what this value counts, where it comes from, and where to look if it is odd, a person seeing the panel for the first time does not have to ask the person who made it. The person who made a dashboard is not always there. And if, while writing a description, you meet a panel that cannot be put in one sentence, that panel does not know what it is looking at either, so it is a candidate for deletion.
What really matters in practice
Once you have built a dashboard, run an incident test. Suppose you were paged at 3 a.m.: can you open this dashboard and say "normal/abnormal" within 30 seconds? If you cannot, it is not that panels are missing but that there are too many.
And write the question it answers right in the dashboard title. "shop-api — are we returning failures to users right now" is better than "shop-api overview." If the title is a question, a panel that does not answer that question stands out.