One Vertical Line Saves Thirty Minutes of Investigation
Goal
You put a deployment and an outage as annotations on a dashboard with a single error rate panel, so that the graph answers not only "when" but also "what happened then." You leave the annotations through the API and read them back to confirm, and you also build a recorder that the deployment pipeline uses to leave them automatically.
Why it matters
Metrics answer only up to the time the value changed. What a person did at that time is an event outside the metrics and is not on the same screen, so the person paged at dawn goes back and forth among deployment records, chat, and the configuration repository to line up the clocks. A large part of the investigation time is that back and forth. An annotation puts that event on the same time axis and removes the back and forth. However, an annotation is a record that cannot be created after the fact, so it must be left at that moment, and for that, a pipeline must leave it, not a human hand. What to write must also be decided in advance — for the next person to be able to roll back right there, the version, the person, and the rollback command must be inside the annotation.
Steps
- Start Grafana with
lab-start-grafanaand create a dashboard with the uidgfd-annot. The panel is a singletimeseriesthat draws the 5xx ratio of shop-api. Then write three lines to/root/gfd-annotations/01-blind.txt— afterquestion=the question this panel answers, afterunanswerable=a question that cannot be answered from this panel alone, and afterreason=why it cannot be answered, in at least 40 characters. - Leave one deployment annotation on this dashboard (
dashboardUIDisgfd-annot) withPOST /api/annotations. The tags must includedeploy, and the body must contain a version in the formversion=vX.Y.Z. Set the time to 30 minutes ago from now (epoch in milliseconds). After reading back withGET /api/annotationsusing theidreturned in the response to confirm, write two lines,id=<그 id>andversion=<적은 버전>, to/root/gfd-annotations/02-annot.txt(the placeholders are that id and the version you wrote). - Leave one region annotation with a start and an end. The tag is
incident, the start is 25 minutes ago from now, and the end is 5 minutes ago from now (so the length is 20 minutes). In the body, write in one line what happened. Then write two lines,id=<그 id>andduration_min=<구간 길이(분)>, to/root/gfd-annotations/03-region.txt(the placeholders are that id and the region length in minutes). - Declare two annotation queries in the dashboard's
annotations.list. One fetches thedeploytag and the other fetches theincidenttag. Both entries must be turned on (enable), have different colors (iconColor), and the target must be the tag-filtering form (target.typeistags). To avoid fetching only annotations that have both tags, put just one tag in each entry. - Create
/root/gfd-annotations/annotate-deploy.sh. It takes a version as its first argument and leaves an annotation on thegfd-annotdashboard, with two tags,deployandauto, and with the three valuesversion=,by=, androllback=in the body. If the environment variableDRY_RUN=1is given, it must not fire and must end after printing only the JSON body to be sent to standard output (the output at that time must be exactly one JSON). Then actually record two different versions with that script. - Write to
/root/gfd-annotations/annotation-fields.txt, one per line, the field names that must be written in a deployment annotation —version,by, androllbackmust be included. Then leave one annotation that keeps to that format on thegfd-annotdashboard. The tags are two,deployandrunbook, and the body must contain all the fields you wrote in the form필드이름=값(the placeholders are the field name and the value). The value ofrollbackmust be a command you can type as is, so it is at least 10 characters. The annotation with therunbooktag must be this one only. - Create
/root/gfd-annotations/07-timeline.tsv. Write all the annotations attached to thegfd-annotdashboard, one per line, in ascending order of time (ascending id if equal), with three tab-separated fields<id> <태그 하나> <요약>(the placeholders are the id, one tag, and the summary). The second field must be one of the tags actually attached to that annotation, and the third field is a summary of at least 4 characters. - Write four lines to
/root/gfd-annotations/08-finding.txt—deploy_id=is the id of the deployment annotation you left in step 2,incident_id=is the id of the region annotation you left in step 3,gap_min=is the difference between the deployment time and the outage start time rounded to an integer in minutes (0 or more), andverdict=is a sentence of at least 60 characters stating the conclusion you drew by linking the two events. Read the two times again from the API to compute.
Notes
- The working directory is
/root/gfd-annotations. The material for the deployment history is/opt/lab/gfd/gfd-annotations/releases.tsv, and the script that made that file is/opt/lab/gfd/gfd-annotations/make.sh. - Start Grafana with
lab-start-grafana(20–40 seconds). It is anonymous Admin, so you can use the API without a token, and you can open the screen through port 3000 of the web preview at the top of the terminal. - The time of an annotation is an epoch integer in milliseconds. If you put it in seconds, it is stamped somewhere in the 1970s and disappears from the screen.
- What cannot be judged in this environment: whether the annotation is really drawn as a vertical line on screen. Panels are drawn by the browser and there is no image renderer plugin. All grading is done only by API responses and the dashboard model — if the query is declared and annotations are retrieved with that tag, the conditions for being drawn are met, but that and "it was drawn" are not the same statement.
- Common mistake: creating without
dashboardUIDso that it becomes an organization-wide annotation. The habit of reading it back withGETright after creating catches this mistake on the spot. - Annotate visualizations · Annotations HTTP API · Dashboard HTTP API · Dashboard JSON model
A graph answers only "when"
Start Grafana with lab-start-grafana and create a dashboard with the uid gfd-annot. The panel is a single timeseries that draws the 5xx ratio of shop-api. Then write three lines to /root/gfd-annotations/01-blind.txt — after question= the question this panel answers, after unanswerable= a question that cannot be answered from this panel alone, and after reason= why it cannot be answered, in at least 40 characters.
Starting Grafana takes 20–40 seconds. Wait until curl -s http://127.0.0.1:3000/api/health gives "database": "ok".
The 5xx ratio is a ratio, not a count. Divide the per-second count of 5xx requests by the per-second count of all requests. Throw it first with promq "<PromQL>" and check that a number comes out.
The reason the panel type must be timeseries becomes clear in the next step. The official documentation says the visualizations that support annotations are Time series, State timeline, and Candlestick — a single-number panel has no place to put events.
Leave the deployment as an annotation and read it back to confirm
Leave one deployment annotation on this dashboard (dashboardUID is gfd-annot) with POST /api/annotations. The tags must include deploy, and the body must contain a version in the form version=vX.Y.Z. Set the time to 30 minutes ago from now (epoch in milliseconds). After reading back with GET /api/annotations using the id returned in the response to confirm, write two lines, id=<그 id> and version=<적은 버전>, to /root/gfd-annotations/02-annot.txt (the placeholders are that id and the version you wrote).
The only required field is text. If you do not write dashboardUID, it becomes an organization-wide annotation and cannot be filtered by this dashboard.
curl -s -X POST -H 'Content-Type: application/json' \
-d '{"dashboardUID":"gfd-annot","time":1700000000000,"tags":["deploy"],"text":"..."}' \
http://127.0.0.1:3000/api/annotations
curl -sG http://127.0.0.1:3000/api/annotations --data-urlencode 'dashboardUID=gfd-annot' | jq
Getting the time unit wrong is the most common mistake. If you put it in seconds, it is stamped somewhere in the 1970s and disappears from the screen. 30 minutes ago is $(( ($(date +%s) - 1800) * 1000 )).
You also use the version in the following steps. The deployment history is in /opt/lab/gfd/gfd-annotations/releases.tsv.
Mark the outage time with a region annotation
Leave one region annotation with a start and an end. The tag is incident, the start is 25 minutes ago from now, and the end is 5 minutes ago from now (so the length is 20 minutes). In the body, write in one line what happened. Then write two lines, id=<그 id> and duration_min=<구간 길이(분)>, to /root/gfd-annotations/03-region.txt (the placeholders are that id and the region length in minutes).
A region annotation is created in the same place as a point annotation. The only difference is that you also send timeEnd. The official documentation says that since Grafana 6.4, a region is represented as a single entry with time and timeEnd.
If you compute the two times separately with date, the seconds drift and the length slips off 20 minutes. Get the reference time once and subtract from it.
NOW=$(date +%s)
echo $(( (NOW - 1500) * 1000 )) $(( (NOW - 300) * 1000 ))
The grader does not simply trust duration_min; it recomputes it from the annotation's two times and cross-checks.
Split by tag and make the dashboard fetch them
Declare two annotation queries in the dashboard's annotations.list. One fetches the deploy tag and the other fetches the incident tag. Both entries must be turned on (enable), have different colors (iconColor), and the target must be the tag-filtering form (target.type is tags). To avoid fetching only annotations that have both tags, put just one tag in each entry.
annotations.list in the dashboard JSON is an array of entries. One entry becomes one toggle at the top of the screen, and its name is the toggle's name.
To point to the built-in annotation data source, set datasource to {"type": "grafana", "uid": "-- Grafana --"}.
curl -s http://127.0.0.1:3000/api/dashboards/uid/gfd-annot \
| jq '.dashboard.annotations.list'
If kinds are mixed, the toggle becomes useless. If you want to see only deployments but outage regions come on along with them, in the end nobody uses the toggle.
Make the pipeline leave it automatically
Create /root/gfd-annotations/annotate-deploy.sh. It takes a version as its first argument and leaves an annotation on the gfd-annot dashboard, with two tags, deploy and auto, and with the three values version=, by=, and rollback= in the body. If the environment variable DRY_RUN=1 is given, it must not fire and must end after printing only the JSON body to be sent to standard output (the output at that time must be exactly one JSON). Then actually record two different versions with that script.
All an automatic recorder has to do is not get skipped even on a busy day. So fix the format and arrange that no person ever calls it by hand.
There are two reasons to have a print mode. A script that runs only inside a pipeline is hard to check when it breaks, and tests must be able to run without side effects. The grader also inspects the body with this mode — and it also checks that the number of annotations does not grow before and after.
If you join strings by hand when building the body, it breaks on quotes. Build it with jq -n --arg. To put the time in as a number, use --argjson.
The deployment history is in /opt/lab/gfd/gfd-annotations/releases.tsv.
Can the next person act right there
Write to /root/gfd-annotations/annotation-fields.txt, one per line, the field names that must be written in a deployment annotation — version, by, and rollback must be included. Then leave one annotation that keeps to that format on the gfd-annot dashboard. The tags are two, deploy and runbook, and the body must contain all the fields you wrote in the form 필드이름=값 (the placeholders are the field name and the value). The value of rollback must be a command you can type as is, so it is at least 10 characters. The annotation with the runbook tag must be this one only.
The content of an annotation is not a matter of taste but a contract. The person who sees that annotation at 3 a.m. must be able to take the next action without opening another window. An annotation with only a commit hash sends that person to the repository.
Writing the rollback command is especially important. Even if the person who deployed is asleep, someone else must be able to roll back, and for that the command must be inside the annotation. The actual commands are in /opt/lab/gfd/gfd-annotations/releases.tsv.
The grader reads the field list you wrote and cross-checks whether the annotation is filled in according to that list.
Reread the annotations in chronological order to make an investigation record
Create /root/gfd-annotations/07-timeline.tsv. Write all the annotations attached to the gfd-annot dashboard, one per line, in ascending order of time (ascending id if equal), with three tab-separated fields <id> <태그 하나> <요약> (the placeholders are the id, one tag, and the summary). The second field must be one of the tags actually attached to that annotation, and the third field is a summary of at least 4 characters.
You have to match the sort key with the grader. If you use sort_by(.time, .id) in jq, the order does not wobble even for annotations stamped in the same millisecond.
curl -sG http://127.0.0.1:3000/api/annotations --data-urlencode 'dashboardUID=gfd-annot' \
| jq -r 'sort_by(.time, .id)[] | [(.id|tostring), .tags[0], .text] | @tsv'
This table is the skeleton of the investigation record. Once you line them up in chronological order, an order like "deployment → error → rollback" catches your eye, and that order is itself the hypothesis about the cause.
Link the deployment and the outage and write a conclusion
Write four lines to /root/gfd-annotations/08-finding.txt — deploy_id= is the id of the deployment annotation you left in step 2, incident_id= is the id of the region annotation you left in step 3, gap_min= is the difference between the deployment time and the outage start time rounded to an integer in minutes (0 or more), and verdict= is a sentence of at least 60 characters stating the conclusion you drew by linking the two events. Read the two times again from the API to compute.
The grader reads the two ids again through the API, computes the gap itself, and cross-checks it against the value you wrote. So a number you estimated by hand will not pass.
curl -sG http://127.0.0.1:3000/api/annotations --data-urlencode 'dashboardUID=gfd-annot' \
| jq -c '.[] | {id, time, timeEnd, tags}'
The conclusion does not have to assert "it was the deployment." A short gap is grounds for suspicion, not evidence. If you also write what the next person should check first, that annotation becomes an investigation record.