TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

A dashboard that had not turned red once in six months

Continue in TT Lab

Goal

After calculating what the default threshold means for our metrics, you pull thresholds from the SLO, use absolute and percentage thresholds separately, make the panel speak in text as well as color, and fix the panel's threshold and an existing alert's threshold so that they pull from one value.

Why it matters

The colors on a dashboard stand in for judgment. So without a basis for the colors, there is no basis for the judgment. Grafana's default threshold is red at 80, and if that value remains on a ratio panel that comes out between 0 and 1, that panel is green no matter what happens. The only threshold you can explain is one computed from a promise made to users — the 30-day availability target sets the error budget, and the error budget sets the threshold. Two more things come with this. A screen that speaks in color alone says nothing to people with color vision deficiency or in black-and-white captures, so value and text must come together, and if the panel's threshold and the alert's threshold are written by hand in different files, they will surely drift apart someday and create a stretch of time where "the screen is red but nobody is paged."

Steps

  1. Start Grafana with lab-start-grafana, create a dashboard with the uid gfd-thresh in /root/gfd-thresholds/dash.json, and upload it. There is one panel, with id 1, type stat, the title 5xx 비율 (기본 문턱 그대로 - 비교용) (the Korean title means "5xx ratio (default threshold as is - for comparison)"), the query sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])), and the unit identifier on the 0..1 ratio side. Leave the threshold as Grafana's default — base green, red at 80, mode absolute. Then write five lines to /root/gfd-thresholds/01-default.txt — now=<지금 이 쿼리가 내는 값>, red_at=<빨강 문턱 값>, max_possible=<이 쿼리가 낼 수 있는 가장 큰 값>, reachable=<빨강에 닿을 수 있으면 yes, 없으면 no>, reason=<왜 그런지 40자 이상> (the placeholders are the value this query produces now, the red threshold value, the largest value this query can produce, yes if it can reach red and no if not, and why, at least 40 characters).
  2. The 30-day availability target is written in /opt/lab/gfd/gfd-thresholds/slo.yml. The error budget is 1 minus that target. Add a stat panel with id 2 to the same dashboard. The title is 5xx 비율 (SLO 문턱) (the Korean title means "5xx ratio (SLO threshold)"), the query is the same as panel 1, and the unit is the same too. The threshold is absolute mode with three steps — the base is green, #EAB839 at the error budget value, and red at double that. Then write five lines to /root/gfd-thresholds/02-slo.txt — budget=<오류 예산>, warn=<노랑 문턱>, crit=<빨강 문턱>, now=<지금 값>, color_now=<지금 값의 색> (the placeholders are the error budget, the yellow threshold, the red threshold, the current value, and the color of the current value). Write the color exactly as the color string written in the panel.
  3. Add a gauge panel with id 3 to the same dashboard. The title is 남은 디스크 (the Korean title means "remaining disk"), the query is node_filesystem_avail_bytes{job="node",mountpoint="/data"}, and the unit is bytes on the 1024 basis. The axis minimum is 0 and the maximum is the data_volume_bytes value in /opt/lab/gfd/gfd-thresholds/slo.yml. The threshold is percentage mode (thresholds.mode is percentage) with three steps — base red, #EAB839 at 20, and green at 40. Then write four lines to /root/gfd-thresholds/03-percentage.txt — avail=<지금 남은 바이트>, max=<최댓값>, pct=<남은 비율을 퍼센트로, 소수 두 자리>, color=<지금 색> (the placeholders are the bytes remaining now, the maximum, the remaining ratio as a percentage to two decimal places, and the current color).
  4. Add three value mappings to panel 2. All are range rules: in order, 정상 (the Korean word for "normal") from 0 to the yellow threshold, 예산 소진 중 (the Korean text means "burning budget") from the yellow threshold to the red threshold, and 사고 (the Korean word for "incident") from the red threshold to 1. Also set this panel's options.textMode to value_and_name so that the number and the name are shown together. Write two lines to /root/gfd-thresholds/04-text.txt — now=<지금 값>, label=<지금 값에 걸리는 글자> (the placeholders are the current value and the text that matches the current value).
  5. Create /root/gfd-thresholds/color.py. When called as python3 color.py <값> [패널 id] (the placeholders are the value and the panel id), it reads that panel's threshold from Grafana and prints which color that value becomes, as one line holding the color string (the default panel id is 2). The rule is the same as Grafana's — after sorting the steps by value, pick the color of the last step whose value is at or below the current value, and treat the first step, whose value is empty, as minus infinity. Then use that tool to create /root/gfd-thresholds/05-boundary.tsv. Six lines without a header, each with two tab-separated fields <값> <색> (the placeholders are the value and the color), with the values in order 0, 0.0049, 0.005, 0.0051, 0.01, and 0.02.
  6. /opt/lab/gfd/gfd-thresholds/alerts.yml is an alert rule already in production. This rule's threshold and the red threshold of panel 2 currently differ. Fix it so that the two numbers are pulled from one place. Write two lines, WARN=<노랑 문턱> and CRIT=<빨강 문턱>, to /root/gfd-thresholds/thresholds.env (the placeholders are the yellow threshold and the red threshold; the values are the ones you computed from the SLO in step 2), change the rule's threshold to the CRIT value of that file, and save it to /root/gfd-thresholds/alerts.yml. Leave the rest of the rule as it is. Then write three lines, without a header, to /root/gfd-thresholds/06-match.tsv, each with two tab-separated fields, writing before, after, and panel in order — the threshold of the original rule, the threshold of the fixed rule, and the panel's red threshold.
  7. Add a timeseries panel with id 4 to the same dashboard. The title is p99 응답 시간 (the Korean title means "p99 response time"), the query is histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))), and the unit is seconds. The threshold is absolute mode with two steps: the base is green, and it is red at the latency_p99_target_seconds value in /opt/lab/gfd/gfd-thresholds/slo.yml. Then write four lines to /root/gfd-thresholds/07-latency.txt — p99_now=<지금 값>, threshold_s=<초로 적은 문턱>, threshold_ms=<같은 문턱을 밀리초로 환산한 값>, color=<지금 색> (the placeholders are the current value, the threshold in seconds, the same threshold converted to milliseconds, and the current color).
  8. /opt/lab/gfd/gfd-thresholds/broken.json is a dashboard pulled from production (uid gfd-thresh-fix). All four panels have a defect in their threshold. Panel 1 has the default threshold left as is and cannot be reached (lower the red to the SLO's crit value), panel 2 is in percentage mode but has no minimum and maximum (put in 0 and data_volume_bytes), panel 3 is a bigger-is-better metric but the step order is reversed (the base must be red and the top green), and panel 4 has a millisecond number as a threshold on a panel in seconds (fix it with latency_p99_target_seconds). Upload the fixed dashboard with the uid gfd-thresh-fix, and write four lines, without a header, to /root/gfd-thresholds/08-report.tsv, each with three tab-separated fields <패널 id> <결함 코드> <무엇이 틀렸었나, 20자 이상이고 숫자를 하나 이상 포함> (the placeholders are the panel id, the defect code, and what was wrong, at least 20 characters and containing at least one number). The defect code is one of unreachable, no-minmax, direction, and scale, and the lines are in panel id order.

Notes

Ask what the default threshold of 80 means for our metric

Start Grafana with lab-start-grafana, create a dashboard with the uid gfd-thresh in /root/gfd-thresholds/dash.json, and upload it. There is one panel, with id 1, type stat, the title 5xx 비율 (기본 문턱 그대로 - 비교용) (the Korean title means "5xx ratio (default threshold as is - for comparison)"), the query sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])), and the unit identifier on the 0..1 ratio side. Leave the threshold as Grafana's default — base green, red at 80, mode absolute. Then write five lines to /root/gfd-thresholds/01-default.txt — now=<지금 이 쿼리가 내는 값>, red_at=<빨강 문턱 값>, max_possible=<이 쿼리가 낼 수 있는 가장 큰 값>, reachable=<빨강에 닿을 수 있으면 yes, 없으면 no>, reason=<왜 그런지 40자 이상> (the placeholders are the value this query produces now, the red threshold value, the largest value this query can produce, yes if it can reach red and no if not, and why, at least 40 characters).

The default threshold is {"mode": "absolute", "steps": [{"color": "green", "value": null}, {"color": "red", "value": 80}]}. The value of the first step being null is the base, and it means minus infinity. This query is a ratio, so its largest value is when the numerator equals the denominator — compare that value with 80.

Pull the threshold from the target

The 30-day availability target is written in /opt/lab/gfd/gfd-thresholds/slo.yml. The error budget is 1 minus that target. Add a stat panel with id 2 to the same dashboard. The title is 5xx 비율 (SLO 문턱) (the Korean title means "5xx ratio (SLO threshold)"), the query is the same as panel 1, and the unit is the same too. The threshold is absolute mode with three steps — the base is green, #EAB839 at the error budget value, and red at double that. Then write five lines to /root/gfd-thresholds/02-slo.txt — budget=<오류 예산>, warn=<노랑 문턱>, crit=<빨강 문턱>, now=<지금 값>, color_now=<지금 값의 색> (the placeholders are the error budget, the yellow threshold, the red threshold, the current value, and the color of the current value). Write the color exactly as the color string written in the panel.

If the target is 0.995, the error budget is 0.005. Do not pick colors by hand; decide them by the rule — Grafana sorts the thresholds by value and then picks the color of the last step whose value is at or below the current value. The boundary is inclusive. You get the current value by throwing it with promq or through the data source proxy.

Absolute and percentage thresholds answer different questions

Add a gauge panel with id 3 to the same dashboard. The title is 남은 디스크 (the Korean title means "remaining disk"), the query is node_filesystem_avail_bytes{job="node",mountpoint="/data"}, and the unit is bytes on the 1024 basis. The axis minimum is 0 and the maximum is the data_volume_bytes value in /opt/lab/gfd/gfd-thresholds/slo.yml. The threshold is percentage mode (thresholds.mode is percentage) with three steps — base red, #EAB839 at 20, and green at 40. Then write four lines to /root/gfd-thresholds/03-percentage.txt — avail=<지금 남은 바이트>, max=<최댓값>, pct=<남은 비율을 퍼센트로, 소수 두 자리>, color=<지금 색> (the placeholders are the bytes remaining now, the maximum, the remaining ratio as a percentage to two decimal places, and the current color).

A percentage threshold compares not the value itself but "where between the minimum and the maximum" against the threshold. So if you do not write the maximum, Grafana guesses from the data on screen, and the color changes every time you change the time range. Remaining capacity is a bigger-is-better value, so the step order is reversed — the base is red.

Do not speak in color alone

Add three value mappings to panel 2. All are range rules: in order, 정상 (the Korean word for "normal") from 0 to the yellow threshold, 예산 소진 중 (the Korean text means "burning budget") from the yellow threshold to the red threshold, and 사고 (the Korean word for "incident") from the red threshold to 1. Also set this panel's options.textMode to value_and_name so that the number and the name are shown together. Write two lines to /root/gfd-thresholds/04-text.txt — now=<지금 값>, label=<지금 값에 걸리는 글자> (the placeholders are the current value and the text that matches the current value).

One value mapping entry has the shape {"type": "range", "options": {"from": 0, "to": 0.005, "result": {"text": "정상", "index": 0}}}. A mapping is scanned from the top and the first rule that matches wins, and both ends are included, so the earlier rule takes the boundary value. The reason for this step is that, for a person with color vision deficiency, red and green are indistinguishable.

Confirm the order and boundaries of thresholds with real values

Create /root/gfd-thresholds/color.py. When called as python3 color.py <값> [패널 id] (the placeholders are the value and the panel id), it reads that panel's threshold from Grafana and prints which color that value becomes, as one line holding the color string (the default panel id is 2). The rule is the same as Grafana's — after sorting the steps by value, pick the color of the last step whose value is at or below the current value, and treat the first step, whose value is empty, as minus infinity. Then use that tool to create /root/gfd-thresholds/05-boundary.tsv. Six lines without a header, each with two tab-separated fields <값> <색> (the placeholders are the value and the color), with the values in order 0, 0.0049, 0.005, 0.0051, 0.01, and 0.02.

You fetch the dashboard with curl -s http://127.0.0.1:3000/api/dashboards/uid/gfd-thresh and pull the steps out of .dashboard.panels[] | select(.id == 2) | .fieldConfig.defaults.thresholds.steps. Use only the standard library (json, sys, urllib.request). Whether the boundary value 0.005 gets the earlier color or the later one is the crux of this step.

Remove the gap where the screen is red but nobody is paged

/opt/lab/gfd/gfd-thresholds/alerts.yml is an alert rule already in production. This rule's threshold and the red threshold of panel 2 currently differ. Fix it so that the two numbers are pulled from one place. Write two lines, WARN=<노랑 문턱> and CRIT=<빨강 문턱>, to /root/gfd-thresholds/thresholds.env (the placeholders are the yellow threshold and the red threshold; the values are the ones you computed from the SLO in step 2), change the rule's threshold to the CRIT value of that file, and save it to /root/gfd-thresholds/alerts.yml. Leave the rest of the rule as it is. Then write three lines, without a header, to /root/gfd-thresholds/06-match.tsv, each with two tab-separated fields, writing before, after, and panel in order — the threshold of the original rule, the threshold of the fixed rule, and the panel's red threshold.

Copy the original to the working directory and just change the number. Check the syntax of the rule file with promtool check rules /root/gfd-thresholds/alerts.yml — the grader uses the same command. How to create a new alert rule is not covered here. The only thing to look at in this step is whether the two numbers are the same.

Application 1 — a latency threshold follows the panel's unit

Add a timeseries panel with id 4 to the same dashboard. The title is p99 응답 시간 (the Korean title means "p99 response time"), the query is histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))), and the unit is seconds. The threshold is absolute mode with two steps: the base is green, and it is red at the latency_p99_target_seconds value in /opt/lab/gfd/gfd-thresholds/slo.yml. Then write four lines to /root/gfd-thresholds/07-latency.txt — p99_now=<지금 값>, threshold_s=<초로 적은 문턱>, threshold_ms=<같은 문턱을 밀리초로 환산한 값>, color=<지금 색> (the placeholders are the current value, the threshold in seconds, the same threshold converted to milliseconds, and the current color).

A threshold must be written in the same unit as the value the panel draws. The value of this panel comes out in seconds, so if the target is 1 second, the threshold is 1, not 1000. If you write it in milliseconds, that panel never turns red — it is the same kind of accident you saw in step 1. You can get the color by giving the panel id to the tool you built in step 5.

Application 2 — fix all the threshold defects of a production dashboard and submit it

/opt/lab/gfd/gfd-thresholds/broken.json is a dashboard pulled from production (uid gfd-thresh-fix). All four panels have a defect in their threshold. Panel 1 has the default threshold left as is and cannot be reached (lower the red to the SLO's crit value), panel 2 is in percentage mode but has no minimum and maximum (put in 0 and data_volume_bytes), panel 3 is a bigger-is-better metric but the step order is reversed (the base must be red and the top green), and panel 4 has a millisecond number as a threshold on a panel in seconds (fix it with latency_p99_target_seconds). Upload the fixed dashboard with the uid gfd-thresh-fix, and write four lines, without a header, to /root/gfd-thresholds/08-report.tsv, each with three tab-separated fields <패널 id> <결함 코드> <무엇이 틀렸었나, 20자 이상이고 숫자를 하나 이상 포함> (the placeholders are the panel id, the defect code, and what was wrong, at least 20 characters and containing at least one number). The defect code is one of unreachable, no-minmax, direction, and scale, and the lines are in panel id order.

Make a copy with cp /opt/lab/gfd/gfd-thresholds/broken.json /root/gfd-thresholds/fixed.json, then fix it and upload it. For some panels you can leave the color strings as they are and touch only the values and order, and for others you have to swap the positions of the colors. After fixing, put in a few values with the tool from step 5 and you can see right away whether the direction is right — just give the panel id as the second argument (though that tool looks at the dashboard whose uid is gfd-thresh).