A Copy Starts Aging Separately That Same Day
In one line
If you copy a dashboard every time a target is added, the copies start aging separately from that day on. A variable turns that copying into one dropdown, and panel repeat unfolds that dropdown into several panels.
Why this was needed
The way a dashboard first breaks is almost always the same. There is one well-made dashboard, and someone who wants to look at another service copies it and changes only the name in the queries. A week later the original gains a panel, and a month later the original's query window changes from [5m] to $__rate_interval. The copy knows about neither.
What makes this state frightening is that the screen looks fine. The copy shows numbers and draws graphs too. It is merely answering a different question from the original, and a person who opened the two dashboards side by side during an incident has to start by finding out why the two numbers differ. The fact that there are as many places to fix as dashboards × panels is also invisible in normal times. You find it out only when you have to fix a query once and end up fixing eight places by hand.
So this module first actually makes the copies. You upload four sets and start by writing down, as a number, "how many places must I fix to correct this query once." With that number, you can feel what the variable removes.
How it works
A template variable goes into the templating.list array of the dashboard JSON. There is the custom type, where you list the values by hand, and the query type, which asks the data source. Something that can grow, like a list of targets, must be query. A hand-written list quietly goes stale the day a fifth target appears.
The moment you let it select several values, one thing changes. To fit a variable with several selected values into a query, you have to turn those values into a single string, and the data source decides how. The official documentation says that Prometheus and InfluxDB use regular expressions, so the values unfold as (값1|값2|값3) (the placeholders are the first, second, and third values) and regular-expression escaping is applied to each value. So the matcher in the query must change too. label="$var" is right only when there is one value, and to accept several values you must write label=~"$var". If you do not change that, it works when you pick one value and becomes an empty graph the moment you pick two. People usually read it as "there's no data" and move on.
Select-all (includeAll) comes with one more setting. If you leave allValue empty, Grafana joins all the values into one long string, and if you write a value (for example .+), that string goes in instead. In a place with hundreds of targets, the latter is much cheaper. However, escaping is not applied, so a person must take responsibility for whether that value is valid in the data source.
Panel repeat starts here. If you put repeat on a panel and specify a multi-value variable, Grafana creates one panel for each selected value. You also decide whether to spread horizontally or vertically (repeatDirection) and how many to put on one row (maxPerRow). Inside a repeated panel, the variable resolves to that panel's single value, so if you put the variable in the title, the title tells you which target a picture is for.
You can also make a variable depend on another variable. If you use the first variable inside the second variable's query, Grafana notices that link and rereads the later value when the earlier value changes. Two things matter here. One is order — variables earlier in templating.list are resolved first, so the dependent one must come later. The other is refresh timing — you choose whether to reread only when the dashboard is opened or also when the time range changes. If you leave a variable whose candidates depend on the time range at the earlier setting, a person who went to look at yesterday sees today's list.
You pay for the convenience. One repeated panel becomes as many panels as there are options, and at least one query goes out for each panel. If there are two repeated panels with four and three options respectively, seven get drawn on screen, and with the fixed panel that makes eight. In an environment with forty targets, the same dashboard becomes eighty panels. So for a dashboard that uses repeat, it is better to set an upper limit on the number of panels that unfold, and above that, first add a variable that narrows the list or split the dashboard.
Finally, I note what you can and cannot verify in this environment. Grading in this lab is done only by the dashboard JSON model and actual query results — whether the variable is declared, what its type is, whether the value list is actually filled from the data source, whether the panel's query uses the variable, and whether that query produces values. Conversely, how many panels were actually drawn on screen cannot be judged. Repeat happens when the browser draws the dashboard, and this Pod has no image renderer plugin. The dashboard JSON that the server returns contains only the single pre-repeat panel, and the variable's options array also comes back empty (null). So the number of options must be counted not from the variable declaration but by asking the data source directly. The only way to see how the repeat actually looks is to check by eye in the web preview.
What it looks like in the field
One team had twelve services and twelve dashboards. An outage required fixing one query, and the person who fixed it did six and went home. The next week, another person looked at one of the remaining six and judged "this one is fine." It was a service running the same code.
There is also what happened from trusting repeat too much. There was a dashboard that set the Pod name as a variable and left select-all as the default, and on the day autoscaling grew the Pods to two hundred, the browser of the person who opened that dashboard froze. It became usable again only after a second variable that narrows the list (namespace) was added in front and the default was changed to one. Repeat is not free; it is a cost multiplied by the number of options.
What you will do in the next lab
First you actually upload four dashboards copied per target and count how many places need fixing. Next, you create a query variable and confirm that its value list is filled from the data, turn on multi-value selection and select-all, and change the query's matcher to a regular expression. You use panel repeat so that as many panels as targets are created, add a second variable that depends on that variable, and decide the order and refresh timing. Finally, you count the number of panels that unfold and write down an upper limit, merge the old dashboard that remains as copies into a single variable, and prove that the merged panel produces the same answer as the original four panels.