TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

An Alert Must Point at a Runbook, Not a Graph

Continue in TT Lab

In one line

Being able to create an alert from a panel and that alert being useful are different matters. What decides usefulness is not the threshold but for and runbook_url.

Why this was needed

The 5xx ratio spikes if even one request fails. In the quiet hours before dawn, a single failure can become several percent. If you page a person the moment it crosses the threshold, it rings several times a day, and then people start ignoring the alerts.

The way alerts die is always this. They are not turned off; they are ignored. And in an alert channel that gets ignored, real outages get buried along with them.

for tackles this problem head on. If you write "if it keeps exceeding for 10 minutes," a momentary spike passes by. If an alert that pages a person has no for, it will surely be ignored someday.

How it works

A Grafana alert rule is a chain of several pieces.

Piece What it does
Query (A) Fetches a value from the data source
Threshold (B) Judges whether that value exceeds the criterion
condition Which piece's result to judge by
for How long it must keep exceeding before paging
annotations What a person reads — a summary and the runbook address

If there is only a query and no threshold, it is not an alert but a metric. That is because when to ring has not been decided.

And you can write all of this in a file.

apiVersion: 1
groups:
  - orgId: 1
    name: shop-api
    folder: Lab
    interval: 1m
    rules:
      - uid: shopapi5xx
        title: ShopApiHighErrorRate
        condition: B
        for: 10m
        annotations:
          summary: 5xx 비율이 10분 넘게 0.5% 를 넘었습니다
          runbook_url: file:///root/graf/runbook.md

An alert rule created in the UI meets the same fate as a dashboard — it remains only inside this Grafana.

If you attach a dashboard link to the alert

The person who got paged clicks the link and a graph appears. The graph says only what is off and does not say what to do. What you need at 3 a.m. is not a picture but a sequence of steps.

A runbook can be short. Three sections are usually enough.

Section What it holds
What broke One or two lines on what this alert means. What is happening to users
What to look at first Commands or queries you can actually type. "Check the status" does not help
How to roll back The action to take before finding the cause. Usually rolling back the latest deployment

It is better to write the runbook before the alert. Then alerts that cannot answer "what does a person do right now when this rings" never get created in the first place.

Common misconceptions

"The more alerts, the safer." Each time one alert is added, the trustworthiness of all the others is chipped away a little. The budget is not the count but attention.

"Just split by severity." Attaching severity: warning and sending it to a channel nobody looks at is not creating an alert; it is creating one more log.

What really matters in practice

There is one thing to ask when creating an alert — "if this rings, what does a person do right now?"

If the answer is "watch it for now," it is not an alert but a dashboard panel. If there is an answer, write that answer in the runbook and make the alert point to that document. Those two lines change the first 5 minutes of the person who got paged.