TT Lab
Get started
Learn Learning paths Courses

Grafana Dashboards

A Graph Only Answers 'When'

Continue in TT Lab

In one line

A graph answers only "when." "Why" is an event outside the metrics, and an annotation puts that event on the same time axis so the next person can find the answer alone.

Why this was needed

A person paged at 3 a.m. opens the dashboard. The error rate turned upward at 02:10. The graph answers up to this point. The next question — "what happened at that time" — the graph cannot answer. So the person opens another window. The record of the deployment pipeline, the scrollback of a chat room, the commit list of a configuration repository. The clocks in the three places differ from each other, and so do the time zones. Of the 20 minutes spent finding the cause, 15 go into this back and forth.

An annotation removes that back and forth. If the pipeline leaves a single line on the graph when a deployment finishes, the next person sees a vertical line right below where the error rate turned. If that line has the version, the person, and how to roll back, they can move on to action instead of investigation.

What matters is that this is a record that cannot be created after the fact. When a deployment finished remains only if you record it at that moment. To reconstruct "was there a deployment then" a month later, you end up digging through pipeline logs, and those logs have usually passed their retention period.

How it works

A Grafana annotation is a single event on the timeline. It can be a point (time only) or a region (time and timeEnd). There are three ways to make one — attaching it directly on a panel, putting it in through the HTTP API, and querying to bring in events that already exist elsewhere.

When creating through the API, the place to use is POST /api/annotations. The official documentation says the only required field is text, dashboardUID and panelId are optional, and if you omit them it becomes an organization-wide annotation. The time is an epoch integer in milliseconds, and when creating a region annotation you also put in timeEnd. After creating it, the response returns an id, and from then on you can read it again with GET /api/annotations. When reading, you can filter by dashboardUID or tags, and if you list tags several times, only those that have all of them (AND) are returned.

On the dashboard side there is the annotation query. If you put an entry in annotations.list of the dashboard JSON and specify a tag, then when that dashboard is opened, Grafana fetches the annotations with that tag and puts them on the panels. Each entry has a name, a color, and an on/off state, so you can switch "deployments only" or "outage regions only" on and off with a toggle at the top of the screen. This is where the value of classifying annotations by tag comes from — if kinds are mixed, the toggle is of no use.

However, annotations are not drawn on every panel. The official documentation says the visualizations that support annotations are Time series, State timeline, and Candlestick. A single-stat panel has no place to put them. For a dashboard built to use annotations, at least one panel must have a time axis.

Annotations a person adds by hand and annotations a pipeline adds automatically differ in nature. Ones added by hand are for pasting in place something found out during an investigation, so the wording is free. For automatic ones, leaving a trace every time is all that matters, so the format must be fixed. If you leave automatic recording to human hands, it gets skipped on a busy day, and that very day is the day you need to find the cause.

So when you build an automatic recorder, you settle two things. One is what to write — the version, the person who deployed (or the pipeline), and a one-line rollback command. Without these three, the person who sees the annotation ends up opening another window. The other is whether you can print it before firing — a script that runs only inside a pipeline is hard to check when it breaks. If you provide a mode that prints the body to be sent as is, a person can inspect it by eye, and tests can run without side effects.

I note what can and cannot be verified in this environment. What can be verified — whether the annotation was actually recorded (read back the id from the POST response with GET and cross-check), whether the time and region are right, whether the tag is attached, whether a query that fetches that tag is declared in the dashboard's annotations.list, and what shape the body the automatic recorder intends to send has. What cannot be verified — whether the annotation is really drawn as a vertical line on screen. Panels are drawn by the browser, and this Pod has no image renderer plugin. So all grading in this lab is done only by the API and the dashboard model, and you must open the picture itself in the web preview and confirm by eye. If the query is declared and annotations are actually retrieved with that tag, the conditions for being drawn are met, but that and "it was drawn" are not the same statement.

What it looks like in the field

The line most often heard on teams that have deployment annotations turned on is "it turned right after the vertical line." That one sentence removes the first 30 minutes of investigation. Conversely, on teams that do not have annotations turned on, for the same outage "isn't it because of the deployment" and "there was no deployment then" go back and forth for 20 minutes.

There are also cases where annotations are turned on but cannot be used. One team's deployment annotation body was just a single commit hash. All the person who received that hash at dawn could do was open the repository, and they ended up waking someone else to ask for the rollback command. The content of an annotation must be decided by whether the next person can act right there.

Where region annotations are especially valuable is in postmortems. If you leave the start and end of an outage as a region, you can recompute the error budget consumption against that region, and next month you do not have to argue over "how many minutes was that outage."

What you will do in the next lab

First you build a dashboard with a single error rate panel and confirm that the cause cannot be known from the graph alone. Next, you record a deployment as an annotation, read it back through the API to cross-check, and mark the outage time with a region annotation that has a start and an end. You split the two by tag and make the dashboard's annotation queries fetch each, and build an automatic recorder for the deployment pipeline to call (including a mode that prints the body to be sent). Finally, you fix in a format what must be written in an annotation for the next person to use it and keep to it, then reread the remaining annotations in chronological order and leave an investigation record and a conclusion.