TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

Write a Dashboard and Build the Tool That Audits It

Continue in TT Lab

Goal

Write down four questions first, and write a health dashboard that answers those questions directly as JSON. Then build a tool that inspects that dashboard, so that a machine asks on your behalf whether each panel has the question it answers written down, whether there are no more than six panels, whether 5xx is viewed as a ratio, and whether latency is viewed as a percentile.

Why it matters

A dashboard grows in only one direction. Anyone can give grounds for adding one because "it would be nice to see this too," but to delete one you have to prove "nobody looks at this," and there is no way to do that. If you write down the question each panel answers, you have grounds to delete a panel that lacks that question, and if you hang the linter on CI, that rule leaves human hands.

The diff of dashboard JSON is hard for a person to read. Coordinates and fields move all over the place, so a review easily just passes. The real purpose of this lab is to create a place where a machine asks on your behalf.

This lab does not start Grafana

What you work with here is the dashboard JSON itself. Starting Grafana and turning it into a screen is done in a later lab of this course. Here you use only files and the linter.

Steps

  1. Write the four questions this dashboard will answer in /root/gfq/01-questions.md. A question is one sentence ending in a question mark, and for each question, on a line starting with metric:, write which metric answers it. It must cover the four golden signals (latency, traffic, errors, saturation).
  2. Write the health dashboard in /root/gfq/dashboard.json. It must have a uid, the title ends in a question mark, and there are 4–6 panels. In each panel's description, write the question that panel answers, ending in a question mark. There must be one panel with 5xx in its title and one with 지연 (the Korean word for "latency") or latency in its title.
  3. Create /root/gfq/lint.py. python3 lint.py <JSON 경로> (the placeholder is the JSON path) prints one line VIOLATION <규칙id> <패널 제목> per violation (the placeholders are the rule id and the panel title) and finally prints violations=<개수> (the placeholder is the count). The first rule is no-description. Save the result of running it on your dashboard to /root/gfq/03-lint-basic.txt.
  4. Add three more rules. too-many-panels (more than 6 panels), error-count-not-ratio (the title has 5xx but the query has no division), latency-not-quantile (the title has 지연 (the Korean word for "latency") or latency but the query has no histogram_quantile). Put in /root/gfq/04-lint-full.txt the results of testing each rule with a file that deliberately violates it.
  5. Deliberately create, in /root/gfq/bad-dashboard.json, a dashboard that violates all four rules, and save the result of running the linter to /root/gfq/05-bad.txt.
  6. Move the diagnosis panels out into /root/gfq/diagnosis.json (three or more panels, a different uid), and make the links of the health dashboard point to that uid. Put the result of checking the panel counts and the link in /root/gfq/06-split.txt.
  7. Add one more rule. description-not-question — if a description exists but does not end in a question mark, it is a violation. Put in /root/gfq/07-lint-e.txt the result of testing with a file whose descriptions are statements.
  8. Write a review in /root/gfq/08-review.md. It needs three sections, ## 30초 시험, ## 지운 패널, and ## CI 에 거는 이유 (the Korean text means "30-second test," "removed panels," and "why put it in CI"), and it must include the distinction among the three kinds: health, diagnosis, and capacity.

Notes

Write the questions first

Go in the order question → panel. If you do it the other way around, it becomes "we have this metric, so let's draw it," and nobody can interpret panels added that way later.

The four golden signals become the four questions as they are.

End each question line with a question mark, and on the line right below it starting with metric: write which metric answers it. The metric line must make clear that errors are a ratio, not a count.

Write the health dashboard as JSON

A health dashboard has 4–6 panels. If it goes over six, diagnosis panels have usually crept in, and then it cannot give an answer within 30 seconds.

In each panel's description, write the question that panel answers. A panel that cannot be put in one sentence does not know what it is looking at either, so it is a candidate for deletion.

Give the title a question form too. If the title is a question, a panel that does not answer that question stands out.

The grader checks whether the query of the panel with 5xx in its title has a division, and whether the panel with 지연 (the Korean word for "latency") or latency in its title uses histogram_quantile.

Build the linter and put in one rule

The output format is the contract. One line VIOLATION <규칙id> <패널 제목> per violation (the placeholders are the rule id and the panel title), and finally one line violations=<개수> (the placeholder is the count). The grader looks for exactly these two formats.

The first rule, no-description, is a violation if a panel's description is missing or empty.

Read with json.load(), and if you accept even the Grafana API shell with doc.get("dashboard", doc), you can use it as is later.

Running it on your dashboard must produce no violations. If it does, go back to step 2 and fill in the descriptions.

Add three rules and test that they catch

Checking only that the linter lets things pass is half the job. You must also feed it files that deliberately violate the rules and see whether it catches them. A linter that always passes is worse than none — it reports that it learned something it never learned, and nobody reports it.

The three rules are these.

Create the test files in a place like /root/gfq/fixtures/ and run the linter on them. You must put that output in the result file so that the rule names remain.

Create a dashboard that violates all four rules

The "wall of graphs" you often see has exactly this shape. Instead of what users experience, process-internal metrics are lined up one after another, 5xx is drawn as a count, latency is an average, and the descriptions are empty.

Put all four into one file.

This file, which you make deliberately, becomes the linter's regression test. Each time you change a rule, you can run it against this.

Move diagnosis panels to another dashboard

The moment you mix diagnosis panels into a health dashboard, that screen cannot give an answer within 30 seconds. You do not delete them — you move them. If you move them and connect them with a link, you can jump over in one step when needed.

The diagnosis dashboard holds panels that look for causes. Things like CPU, GC, connection pool waits, and error rate per handler. The uid must differ from that of the health dashboard.

The links of the health dashboard has a shape like this.

"links": [{"type": "dashboards", "title": "왜 아픈가 — 진단", "url": "/d/<진단 uid>"}]

Put into the result file the output that checks the panel counts and the link of the two dashboards.

Check whether the description is a question too

If a description is a statement, only "what was drawn" remains and "why we look at it" disappears. "Shows the ratio of 5xx requests" is something you can tell by looking at the panel, so it adds nothing.

If you force a question mark, panels that cannot write a question are exposed, and those panels become candidates for deletion.

The rule name is description-not-question. If there is no description at all, that is no-description, so write them separately so that the two rules do not overlap and count twice.

Your dashboard must still have no violations.

Write an incident review

Once you have built a dashboard, run an incident test. Suppose you were paged at 3 a.m.: can you open this dashboard and say within 30 seconds whether things are normal or not? If you cannot, it is not that panels are missing but that there are too many.

Write three sections.

It is also good to note that if the title is a question, a panel that does not answer that question stands out.