TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

How Many Queries a Minute Does This Dashboard Fire?

Continue in TT Lab

Goal

You actually measure the number of queries, series, and points of one dashboard to put a number on it, write a budget as a file and a checker, and then build and submit a version that fits within that budget.

Why it matters

Dashboards look free. Adding one more panel is three clicks, and the cost is incurred on another team's Prometheus. So a dashboard only ever grows in the direction of getting heavier. The arithmetic of putting a number on it is not hard — add the queries attached to each panel, add the variables that ask the server for a list, and that is the number of queries per opening; multiply that by the automatic refresh interval and you get queries per minute. Ten queries at a 10-second refresh is sixty a minute and eighty thousand a day, and they fly out just the same at night when nobody is looking at that screen. A team that has counted this number once and a team that has not treat dashboards differently.

Steps

  1. Start Grafana with lab-start-grafana and upload /opt/lab/gfd/gfd-perf/heavy.json to Grafana as it is, without editing it (the uid is gfd-perf as written in the file, with 8 panels). You can POST it to /api/dashboards/db with curl. You leave this dashboard as the original until the end of this lab — the slimmed version is saved separately under a different uid in the last step.
  2. Download the dashboard uploaded to Grafana and count. Write four lines panels= targets= var_queries= open= to /root/gfd-perf/02-count.txt. panels is the number of panels excluding rows, targets is the number of queries across all panels, var_queries is the number of variables in templating.list whose type is query, and open is targets plus var_queries — the number of queries when the dashboard is opened once.
  3. Actually throw the query attached to each panel and measure the number of series that come back. Pick one past time (at least 5 minutes before now, within 6 hours) and throw all the queries as instant values at that time. Write four lines at= total= top_id= top_series= to /root/gfd-perf/03-series.txt — at is the epoch seconds of the chosen time, total is the sum of the number of series returned by all the dashboard's queries, top_id is the id of the panel that returns the most series, and top_series is the number of series of that panel. Think about how many lines will be drawn in the graph of the heaviest panel.
  4. Pick a past range (at least 2 hours, ending before now) and throw panel 3's p95 query twice over the same range, changing only step — once with step=15 and once with step=300. Write four lines start= end= points_fine= points_coarse= to /root/gfd-perf/04-points.txt. points_fine is the number of points that came back for one series with step 15, and points_coarse is that with step 300.
  5. Read the automatic refresh interval of the dashboard JSON and write four lines refresh= refresh_sec= targets= per_min= to /root/gfd-perf/05-rate.txt. refresh is the string exactly as written in the JSON, refresh_sec is that converted to seconds as an integer, targets is the number of queries counted in step 2, and per_min is targets times 60 / refresh_sec. In this lab's arithmetic, automatic refresh is treated as re-throwing only panel queries and not re-throwing variable queries.
  6. Panel 3's p95 query reads forty-eight bucket series and computes a quantile each time. In /etc/prometheus/rules/gfd-perf.yml, put a group named gfd-perf, an interval of 15 seconds, and one recording rule — record is handler:http_request_duration_seconds:p95 and expr is panel 3's p95 query as it is. After checking with promtool check rules /etc/prometheus/rules/gfd-perf.yml, apply it with curl -X POST http://127.0.0.1:9090/-/reload, wait at least 20 seconds, confirm that promq 'handler:http_request_duration_seconds:p95' returns a value, and then move on to the next step.
  7. Write the budget in five lines to /root/gfd-perf/budget.txt — max_panels=6 max_targets=6 max_var_queries=0 min_refresh_sec=60 max_queries_per_min=6. Then create /root/gfd-perf/budget.py. It takes the budget file path and the dashboard JSON path as arguments and prints violations one per line (starting with B1 to B5), and must end with exit code 1 if there is even one violation. B1 too many panels · B2 too many queries · B3 too many variable queries · B4 refresh is on but the interval is too short · B5 too many queries per minute. Run it on the original /opt/lab/gfd/gfd-perf/heavy.json and confirm that all five violations are caught.
  8. Leave the original gfd-perf as it is and save a new dashboard that fits the budget with the uid gfd-perf-slim. How you slim it is up to you, but these three must be included — remove the variable queries that no panel uses, change the automatic refresh interval to 60 seconds or more, and change the p95 panel's query to the handler:http_request_duration_seconds:p95 you made in step 6. After saving, download that dashboard as it is and save it to /root/gfd-perf/fixed.json (only the .dashboard body), and run the checker from step 7 to confirm 0 violations and exit code 0. Then write to /root/gfd-perf/08-review.md, as five lines B1= through B5=, how you fit each item within the budget, each at least 30 characters.

Notes

Upload the dashboard to put a number on

Start Grafana with lab-start-grafana and upload /opt/lab/gfd/gfd-perf/heavy.json to Grafana as it is, without editing it (the uid is gfd-perf as written in the file, with 8 panels). You can POST it to /api/dashboards/db with curl. You leave this dashboard as the original until the end of this lab — the slimmed version is saved separately under a different uid in the last step.

The save API takes the dashboard body inside the dashboard key and sends overwrite along with it. jq -n --slurpfile is convenient for building that shape from the file. Grafana takes a few tens of seconds to come up, so first check whether /api/health responds.

How many queries fly out when you open it once

Download the dashboard uploaded to Grafana and count. Write four lines panels= targets= var_queries= open= to /root/gfd-perf/02-count.txt. panels is the number of panels excluding rows, targets is the number of queries across all panels, var_queries is the number of variables in templating.list whose type is query, and open is targets plus var_queries — the number of queries when the dashboard is opened once.

A panel folded inside a row is also a panel. To flatten with jq, filter out the ones whose type is row after .panels[] | (., (.panels[]?)). Queries are in each panel's targets array, and one panel can have two or more. For variables, count only the ones that ask the server for a list — a constant list written by hand creates no query.

Which panel is the heaviest

Actually throw the query attached to each panel and measure the number of series that come back. Pick one past time (at least 5 minutes before now, within 6 hours) and throw all the queries as instant values at that time. Write four lines at= total= top_id= top_series= to /root/gfd-perf/03-series.txt — at is the epoch seconds of the chosen time, total is the sum of the number of series returned by all the dashboard's queries, top_id is the id of the panel that returns the most series, and top_series is the number of series of that panel. Think about how many lines will be drawn in the graph of the heaviest panel.

The number of series is the length of data.result in the /api/v1/query response. To pin the time, send time= along with it — that way you get the same answer when you count again later. If you pull the queries out of the dashboard JSON and run them in a loop, you do not have to copy them by hand. If a panel has two queries, the sum of the two is that panel's number of series.

If you change the resolution, how many times more points come back

Pick a past range (at least 2 hours, ending before now) and throw panel 3's p95 query twice over the same range, changing only step — once with step=15 and once with step=300. Write four lines start= end= points_fine= points_coarse= to /root/gfd-perf/04-points.txt. points_fine is the number of points that came back for one series with step 15, and points_coarse is that with step 300.

A range query is /api/v1/query_range and you send start, end, and step along with it. The length of data.result[0].values in the returned JSON is the number of points for one series. If you take the ratio of the two numbers, you can see that resolution multiplies straight into cost. You must pin the range to a past absolute time so that you get the same value when you measure again later.

Multiply by the automatic refresh to get queries per minute

Read the automatic refresh interval of the dashboard JSON and write four lines refresh= refresh_sec= targets= per_min= to /root/gfd-perf/05-rate.txt. refresh is the string exactly as written in the JSON, refresh_sec is that converted to seconds as an integer, targets is the number of queries counted in step 2, and per_min is targets times 60 / refresh_sec. In this lab's arithmetic, automatic refresh is treated as re-throwing only panel queries and not re-throwing variable queries.

The automatic refresh interval is in the top-level refresh key of the dashboard JSON as a string like 10s or 1m. The last character is the unit and the part before it is the number. If you cannot sense why this number matters, multiply it by 24 to get a day's worth — the same number flies out even at night when nobody is looking.

Move the heavy calculation into a recording rule

Panel 3's p95 query reads forty-eight bucket series and computes a quantile each time. In /etc/prometheus/rules/gfd-perf.yml, put a group named gfd-perf, an interval of 15 seconds, and one recording rule — record is handler:http_request_duration_seconds:p95 and expr is panel 3's p95 query as it is. After checking with promtool check rules /etc/prometheus/rules/gfd-perf.yml, apply it with curl -X POST http://127.0.0.1:9090/-/reload, wait at least 20 seconds, confirm that promq 'handler:http_request_duration_seconds:p95' returns a value, and then move on to the next step.

The naming of a recording rule has a convention — written in the form 수준:지표:연산 (the placeholders are level, metric, and operation), with two colons. The group's interval is the calculation interval, and right after applying it, nothing has been calculated yet, so a value appears only after that interval has passed once. You can see whether the rule is loaded with curl -s http://127.0.0.1:9090/api/v1/rules | jq.

Write down the budget and build a tool that checks it

Write the budget in five lines to /root/gfd-perf/budget.txt — max_panels=6 max_targets=6 max_var_queries=0 min_refresh_sec=60 max_queries_per_min=6. Then create /root/gfd-perf/budget.py. It takes the budget file path and the dashboard JSON path as arguments and prints violations one per line (starting with B1 to B5), and must end with exit code 1 if there is even one violation. B1 too many panels · B2 too many queries · B3 too many variable queries · B4 refresh is on but the interval is too short · B5 too many queries per minute. Run it on the original /opt/lab/gfd/gfd-perf/heavy.json and confirm that all five violations are caught.

If you write the budget down as prose, nobody keeps it. Write it as running code. Do not forget that a panel folded inside a row is also a panel, and that the file may be wrapped as {"dashboard": ...}. If refresh is missing or off, there is no automatic refresh, so B4 and B5 must not be triggered.

Slim it down within the budget and submit it as a new version

Leave the original gfd-perf as it is and save a new dashboard that fits the budget with the uid gfd-perf-slim. How you slim it is up to you, but these three must be included — remove the variable queries that no panel uses, change the automatic refresh interval to 60 seconds or more, and change the p95 panel's query to the handler:http_request_duration_seconds:p95 you made in step 6. After saving, download that dashboard as it is and save it to /root/gfd-perf/fixed.json (only the .dashboard body), and run the checker from step 7 to confirm 0 violations and exit code 0. Then write to /root/gfd-perf/08-review.md, as five lines B1= through B5=, how you fit each item within the budget, each at least 30 characters.

When saving under a new uid, you must remove id from the body — if you leave it, Grafana changes the existing dashboard with that id. Reducing the number of queries is not only a matter of deleting panels. See whether there is a panel whose two queries can be merged into a single expression — a panel that receives two counts separately and divides them by eye is one. For a panel that returns forty-eight series, first ask whether there is anything to read from that screen.