What Does It Cost to Open One Dashboard?
In one line
What it costs to open a dashboard is not "the number of panels" but the number of queries times the refresh interval, and nobody counts that value.
Why this was needed
A report that monitoring has gotten slow usually starts with "one query is heavy" and ends with "there are too many dashboards." But there is one number between the two that nobody has counted — how many queries fly out when you open one dashboard once.
Counting it is not hard. Each panel has one or more queries attached, so add all of those, and if any of the variables at the top of the screen ask the server for a list, add those too. With eight panels, ten queries, and two variable queries, it is twelve to open it once. Automatic refresh multiplies on top of that. At a 10-second interval it is sixty times a minute and three thousand six hundred times an hour. One wall-mounted screen throws eighty thousand a day, and throws just the same at night when nobody is looking at that screen.
Worse, this cost does not show on screen at all. Adding one more panel is three clicks, and the cost is incurred on another team's Prometheus. So a dashboard only ever grows in the direction of getting heavier.
How it works
The cost accumulates in three layers.
| Layer | What is counted | Where to read it |
|---|---|---|
| Number of queries | How many on opening, how many per minute | The panels, queries, variables, and refresh in the dashboard JSON |
| Number of series | How many rows one query returns | Actually throw the query and count the result rows |
| Number of points | How many points come per series | Time range divided by resolution |
The first layer is counted just by looking at the JSON. The dashboard JSON model has the panels array, each panel's targets, and refresh, the automatic refresh interval, written as they are (dashboard JSON model documentation). Variables are in templating.list, and the ones that ask the server for a list are those whose type is query.
The second layer is known only by throwing. sum by (handler) (...) returns four series, but if you drop the aggregation and draw rate(http_request_duration_seconds_bucket{...}[5m]) as it is, four handlers times twelve buckets, forty-eight rows, come back. Forty-eight lines are drawn in one panel cell, and there is nothing you can read from that screen. The number of series decides how many points the browser has to draw, and so it is the real reason the screen gets slow.
The third layer is the time range and resolution. A range query takes start, end, and step and returns the span cut at step intervals (Prometheus query API documentation). Viewing two hours at 15-second intervals gives 481 points per series, and at 300-second intervals 25. It is the same question, yet it differs by twenty times. On a 6-hour screen, there is almost never a need for 15-second resolution.
The cheapest way to reduce cost is to precompute answers you already know. Instead of computing a quantile every time from forty-eight bucket series, you have Prometheus do that calculation once at a fixed interval and leave the result as a new time series. That is a recording rule. The dashboard just reads that new time series. The naming has a convention — written in the form 수준:지표:연산 (the placeholders are level, metric, and operation) (recording rule conventions).
What it looks like in the field
The data platform team's Prometheus slowed down every afternoon. The culprit was not the dashboards people open but two screens left up on a meeting room wall. Both refreshed every 5 seconds and had fourteen queries each. Even when nobody was looking, three hundred thirty-six went out per minute. Just changing the refresh to 1 minute and deleting three unused variables cut that load to one twenty-eighth. Not a single panel was deleted.
Another common thing is a variable nobody uses. Once a dropdown at the top of the screen is made, nobody deletes it, yet every time the dashboard is opened, a query flies out to fill its list. It is the same even if no panel references that variable.
What can and cannot be judged in this environment
Measuring cost in milliseconds is meaningless in this Pod. The learner's process and the grader share two CPUs, and other containers run together on the same Mac, so even the same query gives a different time each time you measure. So here you judge only by counts — the number of panels, queries, and variables recounted from the dashboard JSON, the number of series that actually come back when thrown, and the number of points a range query returned. A recording rule is actually applied to Prometheus and checked to see whether its values match the original query. How fast the screen actually appears you ultimately have to see by eye in the browser.
What you will do in the next lab
You receive a dashboard that is over budget and put a number on it. You count panels, queries, and variable queries in the JSON to get "queries per opening," actually throw each query and measure its number of series to find the heaviest panel. You change the time range and resolution and measure the points that come back, and multiply by the refresh interval to get queries per minute. You move the heavy quantile calculation into a recording rule to get the same value more cheaply, write the budget into a file, build a script that checks that budget, and confirm that the slimmed dashboard passes that check.