How Many Queries a Minute Does This Dashboard Fire?
Goal
You actually measure the number of queries, series, and points of one dashboard to put a number on it, write a budget as a file and a checker, and then build and submit a version that fits within that budget.
Why it matters
Dashboards look free. Adding one more panel is three clicks, and the cost is incurred on another team's Prometheus. So a dashboard only ever grows in the direction of getting heavier. The arithmetic of putting a number on it is not hard — add the queries attached to each panel, add the variables that ask the server for a list, and that is the number of queries per opening; multiply that by the automatic refresh interval and you get queries per minute. Ten queries at a 10-second refresh is sixty a minute and eighty thousand a day, and they fly out just the same at night when nobody is looking at that screen. A team that has counted this number once and a team that has not treat dashboards differently.
Steps
- Start Grafana with
lab-start-grafanaand upload/opt/lab/gfd/gfd-perf/heavy.jsonto Grafana as it is, without editing it (the uid isgfd-perfas written in the file, with 8 panels). You can POST it to/api/dashboards/dbwithcurl. You leave this dashboard as the original until the end of this lab — the slimmed version is saved separately under a different uid in the last step. - Download the dashboard uploaded to Grafana and count. Write four lines
panels=targets=var_queries=open=to/root/gfd-perf/02-count.txt.panelsis the number of panels excluding rows,targetsis the number of queries across all panels,var_queriesis the number of variables intemplating.listwhosetypeisquery, andopenistargetsplusvar_queries— the number of queries when the dashboard is opened once. - Actually throw the query attached to each panel and measure the number of series that come back. Pick one past time (at least 5 minutes before now, within 6 hours) and throw all the queries as instant values at that time. Write four lines
at=total=top_id=top_series=to/root/gfd-perf/03-series.txt—atis the epoch seconds of the chosen time,totalis the sum of the number of series returned by all the dashboard's queries,top_idis theidof the panel that returns the most series, andtop_seriesis the number of series of that panel. Think about how many lines will be drawn in the graph of the heaviest panel. - Pick a past range (at least 2 hours, ending before now) and throw panel 3's p95 query twice over the same range, changing only step — once with
step=15and once withstep=300. Write four linesstart=end=points_fine=points_coarse=to/root/gfd-perf/04-points.txt.points_fineis the number of points that came back for one series with step 15, andpoints_coarseis that with step 300. - Read the automatic refresh interval of the dashboard JSON and write four lines
refresh=refresh_sec=targets=per_min=to/root/gfd-perf/05-rate.txt.refreshis the string exactly as written in the JSON,refresh_secis that converted to seconds as an integer,targetsis the number of queries counted in step 2, andper_ministargetstimes60 / refresh_sec. In this lab's arithmetic, automatic refresh is treated as re-throwing only panel queries and not re-throwing variable queries. - Panel 3's p95 query reads forty-eight bucket series and computes a quantile each time. In
/etc/prometheus/rules/gfd-perf.yml, put a group namedgfd-perf, anintervalof 15 seconds, and one recording rule —recordishandler:http_request_duration_seconds:p95andexpris panel 3's p95 query as it is. After checking withpromtool check rules /etc/prometheus/rules/gfd-perf.yml, apply it withcurl -X POST http://127.0.0.1:9090/-/reload, wait at least 20 seconds, confirm thatpromq 'handler:http_request_duration_seconds:p95'returns a value, and then move on to the next step. - Write the budget in five lines to
/root/gfd-perf/budget.txt—max_panels=6max_targets=6max_var_queries=0min_refresh_sec=60max_queries_per_min=6. Then create/root/gfd-perf/budget.py. It takes the budget file path and the dashboard JSON path as arguments and prints violations one per line (starting withB1toB5), and must end with exit code 1 if there is even one violation. B1 too many panels · B2 too many queries · B3 too many variable queries · B4 refresh is on but the interval is too short · B5 too many queries per minute. Run it on the original/opt/lab/gfd/gfd-perf/heavy.jsonand confirm that all five violations are caught. - Leave the original
gfd-perfas it is and save a new dashboard that fits the budget with the uidgfd-perf-slim. How you slim it is up to you, but these three must be included — remove the variable queries that no panel uses, change the automatic refresh interval to 60 seconds or more, and change the p95 panel's query to thehandler:http_request_duration_seconds:p95you made in step 6. After saving, download that dashboard as it is and save it to/root/gfd-perf/fixed.json(only the.dashboardbody), and run the checker from step 7 to confirm 0 violations and exit code 0. Then write to/root/gfd-perf/08-review.md, as five linesB1=throughB5=, how you fit each item within the budget, each at least 30 characters.
Notes
- Start Grafana with
lab-start-grafana(it takes a few tens of seconds). Prometheus is already up. - The original dashboard is in
/opt/lab/gfd/gfd-perf/heavy.json. Do not edit this file; only read it — you use it again in step 7. - The API for saving a dashboard is
POST /api/dashboards/dband the body is{"dashboard": ..., "overwrite": true}. - Rule checking is
promtool check rules <파일>(the placeholder is the file), and applying iscurl -X POST http://127.0.0.1:9090/-/reload. Right after applying, there is no value yet — the first calculation finishes only after the group'sintervalhas passed once. - Cost is measured only by counts. If you measure milliseconds in this Pod, even the same query gives a different value each time — because the learner's process and the grader share two CPUs. How fast the screen actually appears, look at by eye in the web preview.
- Common mistake 1: overwriting the slimmed version onto the same uid. The values counted in earlier steps point to the original, so you leave the original.
- Common mistake 2: not removing
idfrom the body when saving under a new uid. Then an existing dashboard gets changed. - Dashboard JSON model · Prometheus query editor · Prometheus query API · Recording rule configuration · Recording rule naming conventions
Upload the dashboard to put a number on
Start Grafana with lab-start-grafana and upload /opt/lab/gfd/gfd-perf/heavy.json to Grafana as it is, without editing it (the uid is gfd-perf as written in the file, with 8 panels). You can POST it to /api/dashboards/db with curl. You leave this dashboard as the original until the end of this lab — the slimmed version is saved separately under a different uid in the last step.
The save API takes the dashboard body inside the dashboard key and sends overwrite along with it. jq -n --slurpfile is convenient for building that shape from the file. Grafana takes a few tens of seconds to come up, so first check whether /api/health responds.
How many queries fly out when you open it once
Download the dashboard uploaded to Grafana and count. Write four lines panels= targets= var_queries= open= to /root/gfd-perf/02-count.txt. panels is the number of panels excluding rows, targets is the number of queries across all panels, var_queries is the number of variables in templating.list whose type is query, and open is targets plus var_queries — the number of queries when the dashboard is opened once.
A panel folded inside a row is also a panel. To flatten with jq, filter out the ones whose type is row after .panels[] | (., (.panels[]?)). Queries are in each panel's targets array, and one panel can have two or more. For variables, count only the ones that ask the server for a list — a constant list written by hand creates no query.
Which panel is the heaviest
Actually throw the query attached to each panel and measure the number of series that come back. Pick one past time (at least 5 minutes before now, within 6 hours) and throw all the queries as instant values at that time. Write four lines at= total= top_id= top_series= to /root/gfd-perf/03-series.txt — at is the epoch seconds of the chosen time, total is the sum of the number of series returned by all the dashboard's queries, top_id is the id of the panel that returns the most series, and top_series is the number of series of that panel. Think about how many lines will be drawn in the graph of the heaviest panel.
The number of series is the length of data.result in the /api/v1/query response. To pin the time, send time= along with it — that way you get the same answer when you count again later. If you pull the queries out of the dashboard JSON and run them in a loop, you do not have to copy them by hand. If a panel has two queries, the sum of the two is that panel's number of series.
If you change the resolution, how many times more points come back
Pick a past range (at least 2 hours, ending before now) and throw panel 3's p95 query twice over the same range, changing only step — once with step=15 and once with step=300. Write four lines start= end= points_fine= points_coarse= to /root/gfd-perf/04-points.txt. points_fine is the number of points that came back for one series with step 15, and points_coarse is that with step 300.
A range query is /api/v1/query_range and you send start, end, and step along with it. The length of data.result[0].values in the returned JSON is the number of points for one series. If you take the ratio of the two numbers, you can see that resolution multiplies straight into cost. You must pin the range to a past absolute time so that you get the same value when you measure again later.
Multiply by the automatic refresh to get queries per minute
Read the automatic refresh interval of the dashboard JSON and write four lines refresh= refresh_sec= targets= per_min= to /root/gfd-perf/05-rate.txt. refresh is the string exactly as written in the JSON, refresh_sec is that converted to seconds as an integer, targets is the number of queries counted in step 2, and per_min is targets times 60 / refresh_sec. In this lab's arithmetic, automatic refresh is treated as re-throwing only panel queries and not re-throwing variable queries.
The automatic refresh interval is in the top-level refresh key of the dashboard JSON as a string like 10s or 1m. The last character is the unit and the part before it is the number. If you cannot sense why this number matters, multiply it by 24 to get a day's worth — the same number flies out even at night when nobody is looking.
Move the heavy calculation into a recording rule
Panel 3's p95 query reads forty-eight bucket series and computes a quantile each time. In /etc/prometheus/rules/gfd-perf.yml, put a group named gfd-perf, an interval of 15 seconds, and one recording rule — record is handler:http_request_duration_seconds:p95 and expr is panel 3's p95 query as it is. After checking with promtool check rules /etc/prometheus/rules/gfd-perf.yml, apply it with curl -X POST http://127.0.0.1:9090/-/reload, wait at least 20 seconds, confirm that promq 'handler:http_request_duration_seconds:p95' returns a value, and then move on to the next step.
The naming of a recording rule has a convention — written in the form 수준:지표:연산 (the placeholders are level, metric, and operation), with two colons. The group's interval is the calculation interval, and right after applying it, nothing has been calculated yet, so a value appears only after that interval has passed once. You can see whether the rule is loaded with curl -s http://127.0.0.1:9090/api/v1/rules | jq.
Write down the budget and build a tool that checks it
Write the budget in five lines to /root/gfd-perf/budget.txt — max_panels=6 max_targets=6 max_var_queries=0 min_refresh_sec=60 max_queries_per_min=6. Then create /root/gfd-perf/budget.py. It takes the budget file path and the dashboard JSON path as arguments and prints violations one per line (starting with B1 to B5), and must end with exit code 1 if there is even one violation. B1 too many panels · B2 too many queries · B3 too many variable queries · B4 refresh is on but the interval is too short · B5 too many queries per minute. Run it on the original /opt/lab/gfd/gfd-perf/heavy.json and confirm that all five violations are caught.
If you write the budget down as prose, nobody keeps it. Write it as running code. Do not forget that a panel folded inside a row is also a panel, and that the file may be wrapped as {"dashboard": ...}. If refresh is missing or off, there is no automatic refresh, so B4 and B5 must not be triggered.
Slim it down within the budget and submit it as a new version
Leave the original gfd-perf as it is and save a new dashboard that fits the budget with the uid gfd-perf-slim. How you slim it is up to you, but these three must be included — remove the variable queries that no panel uses, change the automatic refresh interval to 60 seconds or more, and change the p95 panel's query to the handler:http_request_duration_seconds:p95 you made in step 6. After saving, download that dashboard as it is and save it to /root/gfd-perf/fixed.json (only the .dashboard body), and run the checker from step 7 to confirm 0 violations and exit code 0. Then write to /root/gfd-perf/08-review.md, as five lines B1= through B5=, how you fit each item within the budget, each at least 30 characters.
When saving under a new uid, you must remove id from the body — if you leave it, Grafana changes the existing dashboard with that id. Reducing the number of queries is not only a matter of deleting panels. See whether there is a panel whose two queries can be merged into a single expression — a panel that receives two counts separately and divides them by eye is one. For a panel that returns forty-eight series, first ask whether there is anything to read from that screen.