A graph is not the data, it is a summary of the data
In one line
A dashboard is not the data but a summary of the data. If there are two ways to summarize, there are two answers, and both are correct. So "it didn't show up on the graph" does not mean "it didn't happen".
Why this matters
In an incident meeting, this exchange happens. "The error rate went up to 28%." "On our dashboard it shows 10%." The two are looking at the same Prometheus, and neither got the query wrong. One just asked with a 5-minute window and the other with a one-hour window.
This mismatch looks trivial but changes decisions. 28% is a number that burns a day's worth of error budget, and 10% is a number that says "let's keep watching". If an alert was set at 15%, it should have fired on one graph and had no reason to fire on the other. In fact, when you dig into "why didn't the alert fire" after an incident, it is common that the culprit is not the rule's threshold but the window length the rule used.
There is a mistaken reading in the opposite direction, too. It is the case where the data exists but the graph is empty. If you use rate(...[15s]) on a metric whose scrape interval is 15 seconds, there is only one sample in the window so a rate of change cannot be computed, and Prometheus returns not an error but an empty result. The screen shows "No data". If you read that as "traffic stopped", you create an outage that does not exist.
How it works
A range query (/api/v1/query_range) takes three things — start, end, and step. Prometheus builds a grid at step intervals from start to end, runs an instant query once at each grid point, and stitches the values together. The line of the graph is not the data but these grid points.
When rate(...[창]) is applied to this, the value at each grid point becomes the average rate of change over the window length immediately before that point. So the window acts as a low-pass filter. If you look at a 20-minute spike with a one-hour window, the value is squashed to about a third, and in exchange the spike smears left and right by the time it is inside the window, so it gets wider. The height and width change but the area underneath is almost conserved — so "how many in total" comes out right and only "how severe was it" comes out wrong.
There is a lower bound for the window as well. The Prometheus query functions documentation and Grafana recommend at least four times the scrape interval. Grafana's $__rate_interval is a variable that automates this rule, calculated as max($__interval + scrape_interval, 4 * scrape_interval). Here $__interval is the panel's time range divided by the panel's pixel width. In other words, if you view the panel wider, the window widens by itself. This is why, if you change the same panel from 12 hours to 30 days, the peak lowers by itself.
The same trap exists in aggregation. If you average ratios with avg, hours with little traffic and hours with much traffic get the same weight. If you want the overall ratio, you must divide the sum by the sum. Quantiles are worse — a percentile is not a value that can be averaged to begin with. For histograms, you must first aggregate the buckets with sum by (le) and then call histogram_quantile, and if you average p99s that have already been computed, you get a number that means nothing.
increase() and rate() extrapolate at both ends of the window. This is because it is rare for samples to sit exactly on the window boundary, and as a result increase() returns a value that should be an integer, like a request count, as a decimal like 20316.33. If you put it on a panel as is, you get the question "what is 0.33 of a request?"
| What you want to ask | What to use | What is commonly used |
|---|---|---|
| How bad it was at its worst | A short window + a fine step | The panel default as is |
| The ratio over the whole period | Sum ÷ sum | The average of ratios |
| Tail latency | histogram_quantile after sum by (le) |
The average of quantiles |
| The exact count | The difference of the counter | The decimal of increase() |
What it looks like in the field
One team used a 30-day dashboard as is in an incident retrospective. That panel's $__rate_interval was over two hours, and the 20-minute incident remained only as a very small bump on the line. The retrospective's conclusion was "the impact was not large". When the same incident was reopened with a 6-hour range, the peak stood at 28%. The only thing that had changed was the time range in the browser address bar.
On another team, a "failure count last night" panel showed 20316.33. The owner wrapped it in round() to get rid of the decimal point, and after that nobody knew that the number was an extrapolated value. During a period where the counter had reset, that panel was quietly showing a wrong value.
What you will do in the next lab
You build a small tool that throws a range query and prints the number of points, the maximum, and the number of values exceeding a threshold, and use it to ask about the same 12 hours in several ways. You shrink the window to 15 seconds and see the graph go empty, widen the window from 1 minute to 1 hour and measure the peak getting lower and wider, and calculate $__rate_interval yourself from the step to reproduce the process by which the panel width decides the window. You compare the average of ratios with sum ÷ sum, confirm why a long-window p99 is not the worst-case p99, and finally fix and submit three broken panels copied from a production dashboard — the grader actually throws the fixed queries and judges by the values.