TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

The same 0.42 meant 420 milliseconds to half the room and 0.42 milliseconds to the rest

Continue in TT Lab

Goal

You upload panels to a real Grafana and put in unit identifiers yourself, confirm by query that the magnitude and unit of the values a query produces match, handle axis ranges and a log axis by hand, and then fix and submit the four unit and axis defects of a dashboard pulled from production.

Why it matters

The numbers on a dashboard can be wrong on screen even when the query is right. A Grafana unit is a display rule, not a conversion rule, so if you attach a millisecond unit to a value that comes out in seconds, the value stays the same and only the name becomes a thousand times smaller. Ratios also have 0..1 and 0..100 as different units, and bytes also differ between the side that shrinks by 1024 and the side that shrinks by 1000. You have to know even that the name shown on screen and the identifier that goes into the JSON are different to be able to leave this work in a file. The axis comes next — if you do not pin the floor to 0, a variation of less than double looks like a cliff, and if you nail down a ceiling, the incident is cut off whole. The purpose of this lab is not to memorize unit names but to become able to look at a single panel and ask "what is this number read as."

Steps

  1. Start Grafana with lab-start-grafana, create a dashboard with the uid gfd-units in /root/gfd-units/dash.json, and upload it to Grafana. There is one panel, with id 1, type timeseries, the title p99 응답 시간 (단위 없음 - 비교용) (the Korean title means "p99 response time (no unit - for comparison)"), and the query histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))). Do not write a unit on this panel (leave it that way in later steps too — it is for comparison). Then write three lines, without a header, to /root/gfd-units/01-readings.tsv, each with two tab-separated fields <가정한 단위 식별자> <그 가정대로면 실제로 몇 초인가> (the placeholders are the assumed unit identifier and how many seconds it actually is under that assumption). The units to assume are, in order, s (seconds), ms (milliseconds), and m (minutes), and write the value converted to seconds to six decimal places.
  2. Add three more panels to the same dashboard. id 2 has the title p99 응답 시간 (the Korean title means "p99 response time") and the query histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))); id 3 has the title 5xx 비율 (the Korean title means "5xx ratio") and the query sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])); id 4 has the title 남은 디스크 (the Korean title means "remaining disk") and the query node_filesystem_avail_bytes{job="node",mountpoint="/data"}. All three have type timeseries, and in each panel's fieldConfig.defaults.unit you write the Grafana unit identifier that fits that value. Latency comes out in seconds, the ratio between 0 and 1, and the disk in bytes. For bytes, use the side that shrinks by 1024 (IEC). Then write three lines, without a header, to /root/gfd-units/02-units.tsv, each with three tab-separated fields <패널 id> <단위 식별자> <이 단위를 고른 이유 15자 이상> (the placeholders are the panel id, the unit identifier, and the reason you chose this unit, at least 15 characters).
  3. Add a timeseries panel with id 5 to the same dashboard. The title is p99 응답 시간 (ms) (the Korean title means "p99 response time (ms)"); use a query that produces the same p99 as a number in milliseconds and use the millisecond unit identifier. Then write two lines, without a header, to /root/gfd-units/03-scale.tsv, each with two tab-separated fields <단위 식별자> <그 패널의 쿼리가 실제로 내는 값> (the placeholders are the unit identifier and the value that panel's query actually produces). The first line is the seconds panel from step 2, the second is this panel, and you write the value to six decimal places.
  4. Add a timeseries panel with id 6 to the same dashboard. The title is 초당 요청 수 (the Korean title means "requests per second"), the query is sum(rate(http_requests_total{job="shop-api"}[5m])), the unit identifier is requests/sec (rps) from the throughput category, and you pin fieldConfig.defaults.min to 0. Conversely, do not put max on the id 2 panel you made in step 2 (if you did, delete it). Then write three lines, without a header, to /root/gfd-units/04-axis.tsv, each with two tab-separated fields, writing rps_min, rps_max, and swing_pct in order. The first two are the smallest and largest values this query produced over the last 12 hours (to three decimal places), and swing_pct is (maximum minus minimum) divided by maximum times 100 (to two decimal places).
  5. Add a timeseries panel with id 7 to the same dashboard. The title is 지연과 오류 비율 (the Korean title means "latency and error ratio"), and there are two queries — histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))) with refId A (legend name p99) and sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])) with refId B (legend name 5xx). Leave the unit of the whole panel as seconds, and in fieldConfig.overrides, select only the series named 5xx, and set its unit to a 0..1 ratio and custom.axisPlacement to right. Then write two lines, without a header, to /root/gfd-units/05-override.tsv, each with three tab-separated fields <범례 이름> <그 계열에 적용되는 단위 식별자> <축 위치> (the placeholders are the legend name, the unit identifier applied to that series, and the axis placement). The axis placement is left or right.
  6. Add a timeseries panel with id 8 to the same dashboard. The title is 핸들러별 초당 요청 수 (the Korean title means "requests per second by handler"), the query is sum by (handler) (rate(http_requests_total{job="shop-api"}[1h])), the unit is requests/sec (rps) from the throughput category, and you set fieldConfig.defaults.custom.scaleDistribution to {"type": "log", "log": 10}. Do not pin the minimum to 0 on this panel. Then write four lines to /root/gfd-units/06-log.txt — top=<가장 큰 핸들러의 값>, bottom=<가장 작은 핸들러의 값> (both to three decimal places; the placeholders are the value of the largest handler and the value of the smallest handler), ratio=<top 나누기 bottom, 소수 두 자리> (the placeholder is top divided by bottom, to two decimal places), and loss=<로그 축으로 바꾸면서 잃는 것, 40자 이상> (the placeholder is what you lose by switching to a log axis, at least 40 characters).
  7. Add two more panels to the same dashboard. id 9 has the title 디스크가 줄어드는 속도 (the Korean title means "rate at which the disk shrinks"), the query - deriv(node_filesystem_avail_bytes{job="node",mountpoint="/data"}[1h]), and the unit is bytes per second on the side that shrinks by 1024. id 10 has the title 디스크가 바닥날 때까지 (the Korean title means "until the disk runs out"), the query is the remaining bytes divided by the shrink rate, and the unit is seconds. Both have type timeseries. Then write three lines, without a header, to /root/gfd-units/07-derived.tsv, each with two tab-separated fields <단위 식별자> <그 쿼리가 내는 값> (the placeholders are the unit identifier and the value that query produces). They are, in order, the remaining bytes, the shrink rate, and the remaining time, and you write the value to three decimal places.
  8. /opt/lab/gfd/gfd-units/broken.json is a dashboard pulled from production (uid gfd-units-fix). All four panels have a defect in their unit or axis. You may leave the queries as they are or change them, but the value shown on screen and the unit must match each other. Upload the fixed dashboard to Grafana with the uid gfd-units-fix. For the axis of panel 4 (requests per second), pin the minimum to 0 and remove the ceiling. Then write four lines, without a header, to /root/gfd-units/08-report.tsv, each with three tab-separated fields <패널 id> <결함 코드> <무엇이 틀렸었나, 20자 이상이고 숫자를 하나 이상 포함> (the placeholders are the panel id, the defect code, and what was wrong, at least 20 characters and containing at least one number). The defect code is one of three — scale (the magnitude of the value and the unit do not match), category (the unit category is wrong), and axis (cut off because of the axis) — and the lines are in panel id order.

Notes

Upload a panel with no unit and read the same number three ways

Start Grafana with lab-start-grafana, create a dashboard with the uid gfd-units in /root/gfd-units/dash.json, and upload it to Grafana. There is one panel, with id 1, type timeseries, the title p99 응답 시간 (단위 없음 - 비교용) (the Korean title means "p99 response time (no unit - for comparison)"), and the query histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))). Do not write a unit on this panel (leave it that way in later steps too — it is for comparison). Then write three lines, without a header, to /root/gfd-units/01-readings.tsv, each with two tab-separated fields <가정한 단위 식별자> <그 가정대로면 실제로 몇 초인가> (the placeholders are the assumed unit identifier and how many seconds it actually is under that assumption). The units to assume are, in order, s (seconds), ms (milliseconds), and m (minutes), and write the value converted to seconds to six decimal places.

You may create the dashboard in the web preview (port 3000) or upload it through the API. The API is curl -s -XPOST -H 'Content-Type: application/json' -d @파일 http://127.0.0.1:3000/api/dashboards/db (the placeholder is the file), and the body you send has the shape {"dashboard": {...}, "overwrite": true}. If you leave the panel's datasource empty, the default data source (Prometheus) is used. The conversion is one multiplication — if you assume minutes, it is that many minutes, so multiply by 60.

Actually put in the unit identifiers

Add three more panels to the same dashboard. id 2 has the title p99 응답 시간 (the Korean title means "p99 response time") and the query histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))); id 3 has the title 5xx 비율 (the Korean title means "5xx ratio") and the query sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])); id 4 has the title 남은 디스크 (the Korean title means "remaining disk") and the query node_filesystem_avail_bytes{job="node",mountpoint="/data"}. All three have type timeseries, and in each panel's fieldConfig.defaults.unit you write the Grafana unit identifier that fits that value. Latency comes out in seconds, the ratio between 0 and 1, and the disk in bytes. For bytes, use the side that shrinks by 1024 (IEC). Then write three lines, without a header, to /root/gfd-units/02-units.tsv, each with three tab-separated fields <패널 id> <단위 식별자> <이 단위를 고른 이유 15자 이상> (the placeholders are the panel id, the unit identifier, and the reason you chose this unit, at least 15 characters).

The name shown on screen and the identifier that goes into the JSON are different. /opt/lab/gfd/gfd-units/unit-picker.md has a table pulled directly from this Pod's Grafana. Ratios split by whether they are between 0 and 1 or between 0 and 100, and for bytes the side that shrinks by 1024 and the side that shrinks by 1000 are different identifiers. If you are curious about the magnitude of a value, throw it first with promq "<쿼리>" (the placeholder is the query).

A unit does not change the value — what do you multiply by to show milliseconds

Add a timeseries panel with id 5 to the same dashboard. The title is p99 응답 시간 (ms) (the Korean title means "p99 response time (ms)"); use a query that produces the same p99 as a number in milliseconds and use the millisecond unit identifier. Then write two lines, without a header, to /root/gfd-units/03-scale.tsv, each with two tab-separated fields <단위 식별자> <그 패널의 쿼리가 실제로 내는 값> (the placeholders are the unit identifier and the value that panel's query actually produces). The first line is the seconds panel from step 2, the second is this panel, and you write the value to six decimal places.

A unit is a display rule, not a conversion rule. If you only attach a millisecond unit to a value that comes out in seconds, the number on screen stays the same and only the name changes — it becomes a panel that reads a thousand times smaller. To make the value milliseconds, you must multiply in the query. The two values must differ by exactly a factor of 1000.

The axis that must start at 0, and the axis that must not get a ceiling

Add a timeseries panel with id 6 to the same dashboard. The title is 초당 요청 수 (the Korean title means "requests per second"), the query is sum(rate(http_requests_total{job="shop-api"}[5m])), the unit identifier is requests/sec (rps) from the throughput category, and you pin fieldConfig.defaults.min to 0. Conversely, do not put max on the id 2 panel you made in step 2 (if you did, delete it). Then write three lines, without a header, to /root/gfd-units/04-axis.tsv, each with two tab-separated fields, writing rps_min, rps_max, and swing_pct in order. The first two are the smallest and largest values this query produced over the last 12 hours (to three decimal places), and swing_pct is (maximum minus minimum) divided by maximum times 100 (to two decimal places).

You get the minimum and maximum over 12 hours with a subquery — write it like min_over_time((<쿼리>)[12h:5m]) (the placeholder is the query). If you leave the axis on automatic, the y-axis starts at the minimum, and a variation of less than double fills the full height of the screen. Conversely, if you nail down a maximum, an incident that happened above it is cut off whole — if you want the line to swing less, use Soft max instead of a hard maximum.

Put two things with different units in one panel

Add a timeseries panel with id 7 to the same dashboard. The title is 지연과 오류 비율 (the Korean title means "latency and error ratio"), and there are two queries — histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))) with refId A (legend name p99) and sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])) with refId B (legend name 5xx). Leave the unit of the whole panel as seconds, and in fieldConfig.overrides, select only the series named 5xx, and set its unit to a 0..1 ratio and custom.axisPlacement to right. Then write two lines, without a header, to /root/gfd-units/05-override.tsv, each with three tab-separated fields <범례 이름> <그 계열에 적용되는 단위 식별자> <축 위치> (the placeholders are the legend name, the unit identifier applied to that series, and the axis placement). The axis placement is left or right.

One override entry has the shape {"matcher": {"id": "byName", "options": "<범례 이름>"}, "properties": [{"id": "unit", "value": "..."}, ...]} (the placeholder is the legend name). You set the legend name with the target's legendFormat. If you overlay them without an override, the two series share one axis, and a ratio of 0.004 becomes a straight line stuck to the floor next to a latency of 0.3.

Where to use a log axis, and what you lose then

Add a timeseries panel with id 8 to the same dashboard. The title is 핸들러별 초당 요청 수 (the Korean title means "requests per second by handler"), the query is sum by (handler) (rate(http_requests_total{job="shop-api"}[1h])), the unit is requests/sec (rps) from the throughput category, and you set fieldConfig.defaults.custom.scaleDistribution to {"type": "log", "log": 10}. Do not pin the minimum to 0 on this panel. Then write four lines to /root/gfd-units/06-log.txt — top=<가장 큰 핸들러의 값>, bottom=<가장 작은 핸들러의 값> (both to three decimal places; the placeholders are the value of the largest handler and the value of the smallest handler), ratio=<top 나누기 bottom, 소수 두 자리> (the placeholder is top divided by bottom, to two decimal places), and loss=<로그 축으로 바꾸면서 잃는 것, 40자 이상> (the placeholder is what you lose by switching to a log axis, at least 40 characters).

You can see the values of the four handlers at once with promq "sum by (handler) (rate(http_requests_total{job="shop-api"}[1h]))", and the largest and smallest values come straight out if you wrap it in max(...) and min(...). On a log axis, 0 cannot be drawn — that is why pinning the minimum to 0 and a log axis cannot be used together. When you write what you lose, think about "what does the same vertical distance come to mean."

Application 1 — what unit goes on the result of a division

Add two more panels to the same dashboard. id 9 has the title 디스크가 줄어드는 속도 (the Korean title means "rate at which the disk shrinks"), the query - deriv(node_filesystem_avail_bytes{job="node",mountpoint="/data"}[1h]), and the unit is bytes per second on the side that shrinks by 1024. id 10 has the title 디스크가 바닥날 때까지 (the Korean title means "until the disk runs out"), the query is the remaining bytes divided by the shrink rate, and the unit is seconds. Both have type timeseries. Then write three lines, without a header, to /root/gfd-units/07-derived.tsv, each with two tab-separated fields <단위 식별자> <그 쿼리가 내는 값> (the placeholders are the unit identifier and the value that query produces). They are, in order, the remaining bytes, the shrink rate, and the remaining time, and you write the value to three decimal places.

If you divide bytes by bytes per second, seconds remain — the unit follows the arithmetic of the query. You get the shrink rate with deriv, and since the value comes out negative, you put a minus sign in front to make it positive. Bytes per second is a unit of a different category from bytes (Data rate, not Data). If you attach the bytes unit as is, the screen says "4 GiB," but what it really means is "4 GiB per second."

Application 2 — fix all the unit and axis defects of a production dashboard and submit it

/opt/lab/gfd/gfd-units/broken.json is a dashboard pulled from production (uid gfd-units-fix). All four panels have a defect in their unit or axis. You may leave the queries as they are or change them, but the value shown on screen and the unit must match each other. Upload the fixed dashboard to Grafana with the uid gfd-units-fix. For the axis of panel 4 (requests per second), pin the minimum to 0 and remove the ceiling. Then write four lines, without a header, to /root/gfd-units/08-report.tsv, each with three tab-separated fields <패널 id> <결함 코드> <무엇이 틀렸었나, 20자 이상이고 숫자를 하나 이상 포함> (the placeholders are the panel id, the defect code, and what was wrong, at least 20 characters and containing at least one number). The defect code is one of three — scale (the magnitude of the value and the unit do not match), category (the unit category is wrong), and axis (cut off because of the axis) — and the lines are in panel id order.

Panel 1 has milliseconds attached to a value that comes out in seconds, and panel 2 has a 0..100 percentage attached to a 0..1 ratio. Each of these has two ways to fix it — match the unit to the value, or multiply the value to fit the unit. Either works. Panel 3 is bytes per second with a bytes unit attached (the category differs). Panel 4 has the right unit but a ceiling on the axis that cuts off the real traffic — measure the 12-hour maximum first. Make a copy with cp /opt/lab/gfd/gfd-units/broken.json /root/gfd-units/fixed.json, then fix it and upload it.