TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

The versions were there, but nobody knew what had changed

Continue in TT Lab

In one line

The heart of managing a dashboard as code is not putting it in a file but making what changed readable and reversible.

Why this was needed

The day after an outage, the dashboard is different. You cannot tell who changed it, and you cannot tell what changed. Grafana keeps versions, but if you download the JSON of two versions and compare them, hundreds of lines come out entirely different. That is because fields that change on every save are mixed in. So there is quite a distance between the fact that "versions are kept" and the fact that "we know what changed."

It is the same when you move it. If you upload a well-made dashboard JSON as is to another environment, all the panels go empty. There is no error on screen; it is just empty. This is because a dashboard refers to a data source by uid, not by name, and auto-generated uids differ per environment. If you do not know this, you lose a day.

How it works

Each time a Grafana dashboard is saved, a new version is created, and the message sent with the save request is attached to that version (dashboard HTTP API). If you leave the message empty, only the time remains in the list. And Grafana overwrites the message of the first save on its own — if you do not check for yourself, you wander for a long time wondering "I wrote a message, so why isn't it kept."

To make two versions readable, you need normalization. What changes on each save are fields like version and id at the top level of the dashboard. If you strip these and emit the keys sorted, from then on the diff becomes something a person can read. A normalization function must have two properties — it always produces the same output for the same input (deterministic), and it makes two versions that differ only in volatile fields the same.

What Why
Remove the top-level version and id They change on every save and bury the diff
Sort the keys If the order wobbles, the same content looks different
Keep the panel id It is the name that refers to a panel, so it carries meaning

The problem of moving is solved with a data source variable. If you create one template variable of the datasource type and have panels point to that variable, the place to fix stays one even when the environment differs (variables documentation). When a query goes out with a uid that does not exist, Grafana returns a 404, but that response is visible only in the browser developer tools, and on the panel it simply appears as an empty graph.

Once you get here, one more property becomes necessary. You must be able to ask for yourself whether the file and the screen are the same. If you only put a dashboard in a repository and nobody checks whether it is actually the same as the screen, the two sides quietly drift apart. Since a normalization function already exists, this check is short — compare the normalized version received from the screen with the normalized file. And that checker also has to be tested in both directions. If you do not confirm that it passes when given the same thing and fails when given a different thing, you end up trusting a checker that passes whatever you give it.

Making things findable is also part of managing as code. When there are dozens of dashboards, which one to open becomes the problem, and what you lean on then are folders and tags. Both are written in the dashboard JSON and the provisioning settings, so once the file becomes the original, even attaching a single tag requires editing the file and applying it. It looks troublesome, but that trouble is exactly the price of leaving "who, when, what, and why."

The last is where to put the original. A dashboard provided through provisioning cannot be saved from the screen or through the API, and an attempt to save is rejected with a 400 (provisioning documentation). It looks inconvenient, but this is the core of this approach — it removes the very path by which the screen and the file could drift apart. From then on, the only way to change the dashboard is to edit the file and have it re-read, and that path comes with a code review.

What it looks like in the field

One team had put its dashboards in a repository, but half a year later the screens and the files were completely different. That was because they had only uploaded them as backups instead of provisioning them from files. Nobody had a reason to edit the files, and nothing stopped editing on screen. The file became the original only after they switched to provisioning.

Another team moved a dashboard to staging and spent two days on "no data is showing." The query was right and Prometheus was fine. The panel's datasource.uid was a value auto-generated in the production environment, and staging had no such uid. After pulling it out into a data source variable, moving became a matter of copying the file.

What can and cannot be judged in this environment

The Grafana in this Pod really runs, but it has no image renderer, so the picture on screen itself cannot be inspected. Instead, the version API, the provisioning state, the save-rejection response, and the dashboard JSON model can all be verified, and this lab judges within that scope. The repository side (review, CI) is not covered because this Pod has no git — instead, we go as far as leaving, as a list, "what must be put in the repository to make a complete set."

What you will do in the next lab

You upload a dashboard, save it once more with a save message, and check what the version list actually keeps. You build a normalization script to extract the difference between two versions in a form a person can read, and pull the data source out into a variable so that it can be moved. Next, you switch the dashboard to be provided from a file and receive for yourself the response that blocks saving. Finally, you attach a single tag by editing the file, and even build a checker that looks for itself at whether the file and the screen are the same.