A dashboard that had not turned red once in six months
Goal
After calculating what the default threshold means for our metrics, you pull thresholds from the SLO, use absolute and percentage thresholds separately, make the panel speak in text as well as color, and fix the panel's threshold and an existing alert's threshold so that they pull from one value.
Why it matters
The colors on a dashboard stand in for judgment. So without a basis for the colors, there is no basis for the judgment. Grafana's default threshold is red at 80, and if that value remains on a ratio panel that comes out between 0 and 1, that panel is green no matter what happens. The only threshold you can explain is one computed from a promise made to users — the 30-day availability target sets the error budget, and the error budget sets the threshold. Two more things come with this. A screen that speaks in color alone says nothing to people with color vision deficiency or in black-and-white captures, so value and text must come together, and if the panel's threshold and the alert's threshold are written by hand in different files, they will surely drift apart someday and create a stretch of time where "the screen is red but nobody is paged."
Steps
- Start Grafana with
lab-start-grafana, create a dashboard with the uidgfd-threshin/root/gfd-thresholds/dash.json, and upload it. There is one panel, withid1, typestat, the title5xx 비율 (기본 문턱 그대로 - 비교용)(the Korean title means "5xx ratio (default threshold as is - for comparison)"), the querysum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])), and the unit identifier on the 0..1 ratio side. Leave the threshold as Grafana's default — base green, red at 80, mode absolute. Then write five lines to/root/gfd-thresholds/01-default.txt—now=<지금 이 쿼리가 내는 값>,red_at=<빨강 문턱 값>,max_possible=<이 쿼리가 낼 수 있는 가장 큰 값>,reachable=<빨강에 닿을 수 있으면 yes, 없으면 no>,reason=<왜 그런지 40자 이상>(the placeholders are the value this query produces now, the red threshold value, the largest value this query can produce, yes if it can reach red and no if not, and why, at least 40 characters). - The 30-day availability target is written in
/opt/lab/gfd/gfd-thresholds/slo.yml. The error budget is 1 minus that target. Add astatpanel withid2to the same dashboard. The title is5xx 비율 (SLO 문턱)(the Korean title means "5xx ratio (SLO threshold)"), the query is the same as panel 1, and the unit is the same too. The threshold is absolute mode with three steps — the base isgreen,#EAB839at the error budget value, andredat double that. Then write five lines to/root/gfd-thresholds/02-slo.txt—budget=<오류 예산>,warn=<노랑 문턱>,crit=<빨강 문턱>,now=<지금 값>,color_now=<지금 값의 색>(the placeholders are the error budget, the yellow threshold, the red threshold, the current value, and the color of the current value). Write the color exactly as the color string written in the panel. - Add a
gaugepanel withid3to the same dashboard. The title is남은 디스크(the Korean title means "remaining disk"), the query isnode_filesystem_avail_bytes{job="node",mountpoint="/data"}, and the unit is bytes on the 1024 basis. The axis minimum is0and the maximum is thedata_volume_bytesvalue in/opt/lab/gfd/gfd-thresholds/slo.yml. The threshold is percentage mode (thresholds.modeispercentage) with three steps — basered,#EAB839at20, andgreenat40. Then write four lines to/root/gfd-thresholds/03-percentage.txt—avail=<지금 남은 바이트>,max=<최댓값>,pct=<남은 비율을 퍼센트로, 소수 두 자리>,color=<지금 색>(the placeholders are the bytes remaining now, the maximum, the remaining ratio as a percentage to two decimal places, and the current color). - Add three value mappings to panel 2. All are
rangerules: in order,정상(the Korean word for "normal") from0to the yellow threshold,예산 소진 중(the Korean text means "burning budget") from the yellow threshold to the red threshold, and사고(the Korean word for "incident") from the red threshold to1. Also set this panel'soptions.textModetovalue_and_nameso that the number and the name are shown together. Write two lines to/root/gfd-thresholds/04-text.txt—now=<지금 값>,label=<지금 값에 걸리는 글자>(the placeholders are the current value and the text that matches the current value). - Create
/root/gfd-thresholds/color.py. When called aspython3 color.py <값> [패널 id](the placeholders are the value and the panel id), it reads that panel's threshold from Grafana and prints which color that value becomes, as one line holding the color string (the default panel id is2). The rule is the same as Grafana's — after sorting the steps by value, pick the color of the last step whose value is at or below the current value, and treat the first step, whose value is empty, as minus infinity. Then use that tool to create/root/gfd-thresholds/05-boundary.tsv. Six lines without a header, each with two tab-separated fields<값> <색>(the placeholders are the value and the color), with the values in order0,0.0049,0.005,0.0051,0.01, and0.02. /opt/lab/gfd/gfd-thresholds/alerts.ymlis an alert rule already in production. This rule's threshold and the red threshold of panel 2 currently differ. Fix it so that the two numbers are pulled from one place. Write two lines,WARN=<노랑 문턱>andCRIT=<빨강 문턱>, to/root/gfd-thresholds/thresholds.env(the placeholders are the yellow threshold and the red threshold; the values are the ones you computed from the SLO in step 2), change the rule's threshold to theCRITvalue of that file, and save it to/root/gfd-thresholds/alerts.yml. Leave the rest of the rule as it is. Then write three lines, without a header, to/root/gfd-thresholds/06-match.tsv, each with two tab-separated fields, writingbefore,after, andpanelin order — the threshold of the original rule, the threshold of the fixed rule, and the panel's red threshold.- Add a
timeseriespanel withid4to the same dashboard. The title isp99 응답 시간(the Korean title means "p99 response time"), the query ishistogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))), and the unit is seconds. The threshold is absolute mode with two steps: the base isgreen, and it isredat thelatency_p99_target_secondsvalue in/opt/lab/gfd/gfd-thresholds/slo.yml. Then write four lines to/root/gfd-thresholds/07-latency.txt—p99_now=<지금 값>,threshold_s=<초로 적은 문턱>,threshold_ms=<같은 문턱을 밀리초로 환산한 값>,color=<지금 색>(the placeholders are the current value, the threshold in seconds, the same threshold converted to milliseconds, and the current color). /opt/lab/gfd/gfd-thresholds/broken.jsonis a dashboard pulled from production (uidgfd-thresh-fix). All four panels have a defect in their threshold. Panel 1 has the default threshold left as is and cannot be reached (lower the red to the SLO'scritvalue), panel 2 is in percentage mode but has no minimum and maximum (put in 0 anddata_volume_bytes), panel 3 is a bigger-is-better metric but the step order is reversed (the base must be red and the top green), and panel 4 has a millisecond number as a threshold on a panel in seconds (fix it withlatency_p99_target_seconds). Upload the fixed dashboard with the uidgfd-thresh-fix, and write four lines, without a header, to/root/gfd-thresholds/08-report.tsv, each with three tab-separated fields<패널 id> <결함 코드> <무엇이 틀렸었나, 20자 이상이고 숫자를 하나 이상 포함>(the placeholders are the panel id, the defect code, and what was wrong, at least 20 characters and containing at least one number). The defect code is one ofunreachable,no-minmax,direction, andscale, and the lines are in panel id order.
Notes
- The working directory is
/root/gfd-thresholds. Start Grafana withlab-start-grafana, and you can also look at it by eye through port 3000 in the web preview. - You may create the dashboard on screen or upload it through the API. The grader looks only at the result uploaded to Grafana. Check the uploaded result with
curl -s http://127.0.0.1:3000/api/dashboards/uid/gfd-thresh | jq '.dashboard.panels'. - Leave a panel's
datasourceempty. The data source uid is generated differently for each Pod. - The materials are in
/opt/lab/gfd/gfd-thresholds—slo.yml, where the targets are written,alerts.yml, already in production, andbroken.json, which you fix in step 8. - This environment cannot judge the color actually painted (there is no image renderer). The grader looks only at the threshold model and the query results, and checks colors by computing exactly the rule Grafana uses. For what has to be seen by eye, open it in the web preview.
- Common mistake: choosing a threshold as a "nice-looking number." The only defensible threshold is one computed from the target.
- Common mistake: leaving the step order as is for a bigger-is-better metric. An empty disk looks green.
- Configure thresholds · Configure standard options · Configure value mappings · Prometheus - Alerting rules · Implementing SLOs (SRE Workbook)
Ask what the default threshold of 80 means for our metric
Start Grafana with lab-start-grafana, create a dashboard with the uid gfd-thresh in /root/gfd-thresholds/dash.json, and upload it. There is one panel, with id 1, type stat, the title 5xx 비율 (기본 문턱 그대로 - 비교용) (the Korean title means "5xx ratio (default threshold as is - for comparison)"), the query sum(rate(http_requests_total{job="shop-api",status=~"5.."}[1h])) / sum(rate(http_requests_total{job="shop-api"}[1h])), and the unit identifier on the 0..1 ratio side. Leave the threshold as Grafana's default — base green, red at 80, mode absolute. Then write five lines to /root/gfd-thresholds/01-default.txt — now=<지금 이 쿼리가 내는 값>, red_at=<빨강 문턱 값>, max_possible=<이 쿼리가 낼 수 있는 가장 큰 값>, reachable=<빨강에 닿을 수 있으면 yes, 없으면 no>, reason=<왜 그런지 40자 이상> (the placeholders are the value this query produces now, the red threshold value, the largest value this query can produce, yes if it can reach red and no if not, and why, at least 40 characters).
The default threshold is {"mode": "absolute", "steps": [{"color": "green", "value": null}, {"color": "red", "value": 80}]}. The value of the first step being null is the base, and it means minus infinity. This query is a ratio, so its largest value is when the numerator equals the denominator — compare that value with 80.
Pull the threshold from the target
The 30-day availability target is written in /opt/lab/gfd/gfd-thresholds/slo.yml. The error budget is 1 minus that target. Add a stat panel with id 2 to the same dashboard. The title is 5xx 비율 (SLO 문턱) (the Korean title means "5xx ratio (SLO threshold)"), the query is the same as panel 1, and the unit is the same too. The threshold is absolute mode with three steps — the base is green, #EAB839 at the error budget value, and red at double that. Then write five lines to /root/gfd-thresholds/02-slo.txt — budget=<오류 예산>, warn=<노랑 문턱>, crit=<빨강 문턱>, now=<지금 값>, color_now=<지금 값의 색> (the placeholders are the error budget, the yellow threshold, the red threshold, the current value, and the color of the current value). Write the color exactly as the color string written in the panel.
If the target is 0.995, the error budget is 0.005. Do not pick colors by hand; decide them by the rule — Grafana sorts the thresholds by value and then picks the color of the last step whose value is at or below the current value. The boundary is inclusive. You get the current value by throwing it with promq or through the data source proxy.
Absolute and percentage thresholds answer different questions
Add a gauge panel with id 3 to the same dashboard. The title is 남은 디스크 (the Korean title means "remaining disk"), the query is node_filesystem_avail_bytes{job="node",mountpoint="/data"}, and the unit is bytes on the 1024 basis. The axis minimum is 0 and the maximum is the data_volume_bytes value in /opt/lab/gfd/gfd-thresholds/slo.yml. The threshold is percentage mode (thresholds.mode is percentage) with three steps — base red, #EAB839 at 20, and green at 40. Then write four lines to /root/gfd-thresholds/03-percentage.txt — avail=<지금 남은 바이트>, max=<최댓값>, pct=<남은 비율을 퍼센트로, 소수 두 자리>, color=<지금 색> (the placeholders are the bytes remaining now, the maximum, the remaining ratio as a percentage to two decimal places, and the current color).
A percentage threshold compares not the value itself but "where between the minimum and the maximum" against the threshold. So if you do not write the maximum, Grafana guesses from the data on screen, and the color changes every time you change the time range. Remaining capacity is a bigger-is-better value, so the step order is reversed — the base is red.
Do not speak in color alone
Add three value mappings to panel 2. All are range rules: in order, 정상 (the Korean word for "normal") from 0 to the yellow threshold, 예산 소진 중 (the Korean text means "burning budget") from the yellow threshold to the red threshold, and 사고 (the Korean word for "incident") from the red threshold to 1. Also set this panel's options.textMode to value_and_name so that the number and the name are shown together. Write two lines to /root/gfd-thresholds/04-text.txt — now=<지금 값>, label=<지금 값에 걸리는 글자> (the placeholders are the current value and the text that matches the current value).
One value mapping entry has the shape {"type": "range", "options": {"from": 0, "to": 0.005, "result": {"text": "정상", "index": 0}}}. A mapping is scanned from the top and the first rule that matches wins, and both ends are included, so the earlier rule takes the boundary value. The reason for this step is that, for a person with color vision deficiency, red and green are indistinguishable.
Confirm the order and boundaries of thresholds with real values
Create /root/gfd-thresholds/color.py. When called as python3 color.py <값> [패널 id] (the placeholders are the value and the panel id), it reads that panel's threshold from Grafana and prints which color that value becomes, as one line holding the color string (the default panel id is 2). The rule is the same as Grafana's — after sorting the steps by value, pick the color of the last step whose value is at or below the current value, and treat the first step, whose value is empty, as minus infinity. Then use that tool to create /root/gfd-thresholds/05-boundary.tsv. Six lines without a header, each with two tab-separated fields <값> <색> (the placeholders are the value and the color), with the values in order 0, 0.0049, 0.005, 0.0051, 0.01, and 0.02.
You fetch the dashboard with curl -s http://127.0.0.1:3000/api/dashboards/uid/gfd-thresh and pull the steps out of .dashboard.panels[] | select(.id == 2) | .fieldConfig.defaults.thresholds.steps. Use only the standard library (json, sys, urllib.request). Whether the boundary value 0.005 gets the earlier color or the later one is the crux of this step.
Remove the gap where the screen is red but nobody is paged
/opt/lab/gfd/gfd-thresholds/alerts.yml is an alert rule already in production. This rule's threshold and the red threshold of panel 2 currently differ. Fix it so that the two numbers are pulled from one place. Write two lines, WARN=<노랑 문턱> and CRIT=<빨강 문턱>, to /root/gfd-thresholds/thresholds.env (the placeholders are the yellow threshold and the red threshold; the values are the ones you computed from the SLO in step 2), change the rule's threshold to the CRIT value of that file, and save it to /root/gfd-thresholds/alerts.yml. Leave the rest of the rule as it is. Then write three lines, without a header, to /root/gfd-thresholds/06-match.tsv, each with two tab-separated fields, writing before, after, and panel in order — the threshold of the original rule, the threshold of the fixed rule, and the panel's red threshold.
Copy the original to the working directory and just change the number. Check the syntax of the rule file with promtool check rules /root/gfd-thresholds/alerts.yml — the grader uses the same command. How to create a new alert rule is not covered here. The only thing to look at in this step is whether the two numbers are the same.
Application 1 — a latency threshold follows the panel's unit
Add a timeseries panel with id 4 to the same dashboard. The title is p99 응답 시간 (the Korean title means "p99 response time"), the query is histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{job="shop-api"}[6h]))), and the unit is seconds. The threshold is absolute mode with two steps: the base is green, and it is red at the latency_p99_target_seconds value in /opt/lab/gfd/gfd-thresholds/slo.yml. Then write four lines to /root/gfd-thresholds/07-latency.txt — p99_now=<지금 값>, threshold_s=<초로 적은 문턱>, threshold_ms=<같은 문턱을 밀리초로 환산한 값>, color=<지금 색> (the placeholders are the current value, the threshold in seconds, the same threshold converted to milliseconds, and the current color).
A threshold must be written in the same unit as the value the panel draws. The value of this panel comes out in seconds, so if the target is 1 second, the threshold is 1, not 1000. If you write it in milliseconds, that panel never turns red — it is the same kind of accident you saw in step 1. You can get the color by giving the panel id to the tool you built in step 5.
Application 2 — fix all the threshold defects of a production dashboard and submit it
/opt/lab/gfd/gfd-thresholds/broken.json is a dashboard pulled from production (uid gfd-thresh-fix). All four panels have a defect in their threshold. Panel 1 has the default threshold left as is and cannot be reached (lower the red to the SLO's crit value), panel 2 is in percentage mode but has no minimum and maximum (put in 0 and data_volume_bytes), panel 3 is a bigger-is-better metric but the step order is reversed (the base must be red and the top green), and panel 4 has a millisecond number as a threshold on a panel in seconds (fix it with latency_p99_target_seconds). Upload the fixed dashboard with the uid gfd-thresh-fix, and write four lines, without a header, to /root/gfd-thresholds/08-report.tsv, each with three tab-separated fields <패널 id> <결함 코드> <무엇이 틀렸었나, 20자 이상이고 숫자를 하나 이상 포함> (the placeholders are the panel id, the defect code, and what was wrong, at least 20 characters and containing at least one number). The defect code is one of unreachable, no-minmax, direction, and scale, and the lines are in panel id order.
Make a copy with cp /opt/lab/gfd/gfd-thresholds/broken.json /root/gfd-thresholds/fixed.json, then fix it and upload it. For some panels you can leave the color strings as they are and touch only the values and order, and for others you have to swap the positions of the colors. After fixing, put in a few values with the tool from step 5 and you can see right away whether the direction is right — just give the panel id as the second argument (though that tool looks at the dashboard whose uid is gfd-thresh).