Build a Dashboard That Answers One Question
Goal
You start a real Grafana, attach a real Prometheus, and build a dashboard that answers one question from scratch. When it is finished, you solidify it into provisioning files so that even if you delete this Grafana, the same screen can be set up again.
This is the first lab with a screen. If you start Grafana at http://127.0.0.1:3000, you can open the real Grafana screen with the web preview button at the top of the terminal. You may create panels by clicking on that screen or upload them through the API — the grader does not ask which way you built them and looks only at the result uploaded to Grafana.
Why it matters
A dashboard is not a picture but a tool that answers a question. That is why this lab goes not in the order of drawing panels but in the order question → query → panel → alert → file.
The last step is especially important. A dashboard created by clicking exists only inside Grafana's database, so when that database disappears it disappears with it, and no record remains of who changed what and when. In practice, failing to find "that dashboard from back then" is almost always for this reason.
This Pod contains 12 hours of metrics for a fictional service called shop-api. http_requests_total has four handler labels (/api/orders, /api/search, /api/users, /healthz), and response time comes in as the http_request_duration_seconds_bucket histogram.
Steps
- Declare the data source in
/root/graf/provisioning/datasources/prometheus.ymland start Grafana athttp://127.0.0.1:3000. The data source's uid islabpromand its address ishttp://127.0.0.1:9090. - Write the question this dashboard will answer in one sentence in
/root/graf/02-question.md, and the PromQL that answers that question on one line in/root/graf/02-question.promql. - Create and upload a dashboard with one panel containing that query. The uid is
shop-api. - Add a p95 response time panel and a current request rate panel. The type follows the shape of the question.
- Add a
handlertemplate variable and make every panel narrow by that variable. - Write a runbook in
/root/graf/runbook.md— what broke / what to look at first / how to roll back. - Create an alert rule in
/root/graf/provisioning/alerting/shop-api.ymland make it point to the runbook. - Export the dashboard as JSON and place it in
/root/graf/dashboards/shop-api.json, and make Grafana watch that directory with a provider file.
Notes
lab-statusshows what is running right now, andpromq "<PromQL>"lets you throw a query directly.- This lab does not use
lab-start-grafana. That helper starts Grafana with the data source already attached, so you could not try step 1. If it is already on, take it down withpkill -x grafanaand start over. - Provisioning configuration files are read at startup. If you edited a file and the screen did not change, it is usually because you did not start it again.
- Starting Grafana takes 20–40 seconds. Wait until
curl -s http://127.0.0.1:3000/api/healthgives"database": "ok". - When you start Grafana again, wait until the previous process lets go of 3000. If you start a new one right after
pkill -x grafana, the new process cannot grab the port and dies.
Start Grafana and attach the data source through a file
Declare a prometheus data source with the uid labprom in /root/graf/provisioning/datasources/prometheus.yml, and start Grafana at http://127.0.0.1:3000 with that directory set as the provisioning path.
If you attach the data source by clicking in the UI, it remains only in Grafana's DB. The grader checks whether readOnly in the API response is true — that is the mark that it "came from a file."
mkdir -p /root/graf/provisioning/datasources /root/graf/provisioning/dashboards \
/root/graf/provisioning/alerting /root/graf/dashboards \
/tmp/gf/data /tmp/gf/logs /tmp/gf/plugins
GF_PATHS_DATA=/tmp/gf/data GF_PATHS_LOGS=/tmp/gf/logs GF_PATHS_PLUGINS=/tmp/gf/plugins \
GF_PATHS_PROVISIONING=/root/graf/provisioning \
GF_SERVER_HTTP_PORT=3000 \
GF_AUTH_ANONYMOUS_ENABLED=true GF_AUTH_ANONYMOUS_ORG_ROLE=Admin \
GF_PLUGINS_PREINSTALL_DISABLED=true \
setsid nohup grafana server --homepath /opt/grafana >/var/log/grafana.log 2>&1 </dev/null &
curl -s http://127.0.0.1:3000/api/health # "database": "ok" 가 나올 때까지 20~40초
curl -s http://127.0.0.1:3000/api/datasources/uid/labprom/health
The reasons for the three groups of environment variables are these. Redirecting the data, log, and plugin paths to /tmp is because this Pod has no capabilities and may not be able to write to the default paths; turning on anonymous access is so that you do not meet a login screen first when opening through the web preview; and turning off plugin pre-installation is because Grafana 11.4 tries to download grafana-lokiexplore-app from the internet on every startup, but this Pod cannot reach the outside.
Write down the question this dashboard will answer first
In /root/graf/02-question.md, write the question this dashboard will answer as one sentence ending in a question mark, and in /root/graf/02-question.promql, write the PromQL that answers that question. The question is "What percent of shop-api's responses are 5xx?"
It is a ratio, not a count. When traffic doubles, the 5xx count doubles too, but the probability of users experiencing failure stays the same.
Divide the per-second count of 5xx requests by the per-second count of all requests. Use the status label and rate(). Set the window to [5m].
promq 'sum(rate(http_requests_total[5m]))'
promq 'sum by (status) (rate(http_requests_total[5m]))'
promq "$(cat /root/graf/02-question.promql)"
The grader looks at the value produced by actually running the query you wrote through the Grafana data source. Even if you write it differently, you pass if the number is right.
Create a panel that answers the question
Create a dashboard with the uid shop-api and put in one timeseries panel that draws the query from step 2. Name the dashboard title so that the question it answers can be seen.
You may open Grafana in the web preview, create a new dashboard, and add a panel (in that case, set the uid to shop-api in the dashboard settings), or upload through the API as below.
curl -s -XPOST -H 'Content-Type: application/json' \
-d @/tmp/dash.json http://127.0.0.1:3000/api/dashboards/db
/tmp/dash.json has the shape {"overwrite": true, "dashboard": { ... }}.
The type must be timeseries. The answer you need in an incident is not "what percent now" but "since when did it rise," and you cannot see that with a type that has no time axis.
Check the uploaded result like this.
curl -s http://127.0.0.1:3000/api/dashboards/uid/shop-api | jq '.dashboard.panels'
Match the panel type to the shape of the question
Add a p95 response time panel and a current request rate panel to the same dashboard to make three panels. p95 is timeseries and the current request rate is stat.
p95 is computed from the histogram. le is the bucket boundary, so you must keep only that and sum the rest.
promq 'histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))'
promq 'sum(rate(http_requests_total[5m]))'
Do not use gauge for the current request rate. A gauge is a tool that shows where between 0 and the maximum, but requests per second has no maximum. A single number at this very moment is stat.
Conversely, if you put p95 in stat or gauge, you cannot see "since when did it get slow." For a quantile, how it moved over time is everything.
Make one dashboard view four handlers
Add a query-type template variable named handler, and make the queries of all three panels narrow by that variable.
A custom variable where you list the values by hand goes stale the day a handler is added. Read it from the data.
label_values(http_requests_total, handler)
And a variable does nothing just by being declared. The panel queries must be narrowed like this.
sum(rate(http_requests_total{handler="$handler"}[5m]))
The grader swaps the variable for actual values and throws the query. If the label name or the quotes are wrong, an empty result comes back and you fail on the spot.
Write the runbook before the alert
Write a runbook in /root/graf/runbook.md. It must have three sections, ## 무엇이 깨졌나, ## 먼저 볼 것, and ## 되돌리는 법 (the Korean text means "what broke," "what to look at first," and "how to roll back"), and the "what to look at first" section must include a http_requests_total query you can actually throw.
There is a reason for this order. If you create the alert first, it becomes "we'll think about it when it rings," and that document never gets written. If you write the runbook first, alerts that cannot answer "what does a person do right now when this rings" never get created in the first place.
"What to look at first" must be a command, not a sentence. "Check the status" is no help at 3 a.m. Write down, as is, a query that counts which handler is failing.
Create the alert and make it point to the runbook
In /root/graf/provisioning/alerting/shop-api.yml, create an alert rule for the 5xx ratio. It must have a threshold condition, for must be 5 minutes or more, and annotations.runbook_url must point to the runbook from step 6. After writing the file, start Grafana again.
for is the heart of this step. 5xx spikes briefly if even one request fails, and if you wake a person every time, that person will ignore alerts from then on. Alerts die not by being turned off but by being ignored.
You cannot use dashboard variables like $handler in an alert query. An alert is evaluated without a screen, so there is no dropdown to resolve the variable.
Provisioning configuration is read at startup. If you only write the file, nothing happens.
pkill -x grafana; sleep 3
# (1단계와 같은 환경변수로 다시 띄운다)
curl -s http://127.0.0.1:3000/api/v1/provisioning/alert-rules | jq '.[].title'
There is also a way to have it re-read with an admin account (an anonymous request gets a 403).
curl -XPOST -u admin:admin http://127.0.0.1:3000/api/admin/provisioning/alerting/reload
Solidify the dashboard into a file
Export the dashboard now showing on screen as JSON and save it to /root/graf/dashboards/shop-api.json, make Grafana watch that directory with /root/graf/provisioning/dashboards/lab.yml, and start it again. The grader checks whether meta.provisioned is true.
Dashboard JSON and the provider file are different things. The provider file holds not a dashboard but an instruction to watch a certain directory.
curl -s http://127.0.0.1:3000/api/dashboards/uid/shop-api \
| jq '.dashboard' > /root/graf/dashboards/shop-api.json
pkill -x grafana; sleep 3
# (1단계와 같은 환경변수로 다시 띄운다)
curl -s http://127.0.0.1:3000/api/dashboards/uid/shop-api | jq '.meta.provisioned'
If meta.provisioned is true, what is on the screen now came from a file. From then on, even if you try to overwrite it through the API, Grafana refuses with Cannot save provisioned dashboard — this blocks the path by which the screen and the file could drift apart.