TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

A number with no unit is a number the reader makes up

Continue in TT Lab

In one line

A number with no unit gets a unit invented by whoever reads it. And the invented unit is usually the one that favors that person.

Why this was needed

In an outage meeting, someone pulls up a dashboard and says, "Latency is 0.42." Half the room heard 420 milliseconds and half heard 0.42 milliseconds. The two numbers differ by a factor of a thousand, one is an incident and the other is something to brag about. The conversation went on for another 20 minutes or so, and ended only when someone opened the query.

This kind of mismatch does not happen because people are careless. It happens because the screen shows only the number. Grafana has a feature that attaches a unit to a value, but that feature works only when we tell it what the value is. If you say nothing, Grafana draws the number as it is, and interpretation is left to whoever is looking.

The worse case is when a unit is written but written wrongly. Without a unit, people at least doubt. With a unit attached, nobody doubts. If you attach a millisecond unit to a query that comes out in seconds, the screen says "0.42 ms" with confidence, and the person who sees it believes the service is very fast.

How it works

A Grafana unit is a display rule, not a conversion rule. If you write an identifier in the panel's fieldConfig.defaults.unit, Grafana turns the number into a string with the formatting function tied to that identifier. The value itself is not touched. So to make a value that comes out in seconds show as milliseconds, it is not enough to change the unit — you have to multiply the query by 1000.

The identifier differs from the name shown on screen. The dropdown says "Percent (0.0-1.0)," but the value that goes into the JSON is percentunit. "bytes(IEC)" is bytes and "bytes(SI)" is decbytes. The former divides by 1024 and shrinks to GiB, and the latter divides by 1000 and shrinks to GB. The same 4294967296 shows as 4 GiB or as 4.29 GB. Which is right depends on what the side that produced the number was counting.

Ratios come in two sets as well. For values between 0 and 1, use percentunit, and for values between 0 and 100, use percent. An error ratio built with PromQL is almost always the former, yet it is common for the latter to be attached to the panel. Then an error rate of 1.6% shows on screen as 0.016%. This is how you end up in a state where the alert fires but the dashboard looks calm.

The axis is the place of the second lie, after the unit. By default, Grafana sets the y-axis automatically to fit the data. If requests per second wobble between 34 and 75, the axis is set to 34 to 75 too, and within it the line fills the full height of the screen as it swings. In reality it is a difference of a little over double, but the screen looks like a cliff. So it is more honest to pin the minimum of a panel that counts quantities to 0. Min and Max in the standard options documentation are the place for that.

Conversely, you must not pin the maximum carelessly. A request-count panel with the maximum fixed at 20 shows a flat ceiling cut off at 20 even on a day when traffic rose to 75. Incidents always happen above that ceiling, but nothing remains on screen. If you dislike the line looking flat because the variation is small, use the Soft min and Soft max of the time series panel instead of a hard maximum. If the data goes beyond that range, the axis widens to follow so that nothing is cut off.

Sometimes you have to draw two things with different units in one panel. A panel that overlays latency and error ratio and asks "is the time it got slower the same as the time failures increased" is an example. In that case, you set a single unit for the whole panel and give the unit and axis placement separately only to the remaining series, using overrides. The time series panel documentation says that when there are two or more units, the first unit uses the left axis and the units after it use the right axis. If you just overlay them without an override, the two series share one axis, and a ratio of 0.004 becomes a straight line stuck to the floor next to a latency of 0.3.

A log axis is used when you put series of very different magnitudes on one screen. If you draw a handler at 42 requests per second and one at 6 on the same linear axis, the changes of the smaller one cannot be seen. On a log axis, the same vertical distance means the same multiple, so both become readable. But there is something you lose. The distance between ticks is no longer a difference in quantity, so you cannot add up areas by eye, zero and negative numbers cannot be drawn at all, and an event that doubled and one that rose tenfold look like similar height differences. So it does not suit a panel that asks "how much did it grow."

There is one thing this lab cannot judge — the picture actually drawn. The Grafana in this Pod has no image renderer plugin, so panels cannot be rendered as images. So the grader looks only at the dashboard JSON model and the query results. It can verify what the unit identifier is, what the minimum of the axis is, and what values the query actually produces, but it cannot verify whether that combination is readable on screen. Problems like axis names overlapping and being cut off, or the legend covering the graph, you have to see with your own eyes by opening port 3000 in the web preview. That the model is right and that the screen is readable are different matters.

What it looks like in the field

One team had an SI unit attached to its memory usage panel. The container limit was 4 GiB, but the screen showed 4.29 GB, and people read it as "there is still room." That night, the Pod hit its limit and died. The value was never wrong even once; only the name was wrong.

On another team, the error rate panel had been attached as percent for months. Even on a day when the real error rate rose to 2%, the screen showed 0.02%. The alert fired normally, but the person on duty looked at the dashboard, judged that "the alert seems to have fired in error," and ignored it. What the incident retrospective talked about longest was not the alert rule but that single character of unit.

What you will do in the next lab

You start a real Grafana, upload one panel with no unit, and write down for yourself in how many ways that number can be read. Next, you put Grafana's unit identifiers on seconds, ratio, and bytes panels, throw queries to confirm that the magnitude and unit of the values match, and confirm by hand what you must multiply by for it to show as milliseconds. You separate the places where the axis minimum must be pinned to 0 from the places where the maximum must not be pinned, put two series with different units into one panel using an override, and after using a log axis write down what you lost. At the end you receive a dashboard pulled from production and submit it with all four unit and axis defects fixed — the grader asks Grafana for the fixed dashboard and actually throws the panels' queries to check that the values and units match.