The Day Every Dashboard Went Blank
In one line
Envoy statistic names are a structure in which one string carries several meanings (cluster.<이름>.upstream_rq_total, where the placeholder is the cluster name). When exported to Prometheus, that string is split into a metric name and labels, and you can add the splitting rules yourself. And the name is the interface, so changing even one word leaves the dashboards that used it empty.
Why this was needed
You can tell what is happening in a proxy from logs too, but logs pile up one line per request, so they are expensive for asking "how is everything right now". Statistics are the opposite — they do not grow with the number of requests, and they hold aggregated numbers from the start.
The problem is that there are more than five hundred of those numbers. That is so even in the default configuration, and as clusters grow, dozens more are attached for each cluster. So working with statistics becomes not "what exists" but "how do I find what I need, and how do I trim what I do not need".
How it works
Three branches of names. The first word of the name says what the number is about.
| Prefix | What the number is about | Representative |
|---|---|---|
cluster.<이름>. (the placeholder is the cluster name) |
The upstream side | upstream_rq_total, membership_healthy |
http.<stat_prefix>. |
The requests handled by that HTTP connection manager | downstream_rq_5xx |
listener.<주소>. (the placeholder is the address) |
The socket side | downstream_cx_total |
One and the same request is counted in three places separately. So when the three numbers disagree, that difference is the clue — if a connection was caught at the listener but the request count on the http. side is low, a connection was made and no request was sent, and if 5xx on the http. side increased but the request count on the cluster. side stayed the same, it did not even reach the upstream.
The kinds differ. A counter is cumulative and only grows, a gauge is the current state, and a histogram is a distribution. /reset_counters on the admin port, as its name says, resets only counters. A gauge is a current state, such as "how many servers are healthy right now", so there is nothing to reset.
How to pull them out. You can attach ?filter=<정규식> (the placeholder is a regular expression) and ?format=json to /stats, and the two can be used together. When scraping with Prometheus you use /stats/prometheus, and here the name is split. cluster.good.upstream_rq_total becomes one metric name and a label carrying the cluster name. Envoy holds several splitting rules by default, and you can add more with stats_config.stats_tags — if you put a team name as a prefix on cluster names, you can pull out just that as a label and group by team on the dashboard.
How to trim. With stats_config.stats_matcher you set an inclusion list or an exclusion list. One important property here — an excluded statistic is not hidden from the display, it is not recorded at all. Even if you bring it back later when you need it, there are no values from the time in between.
What it looks like in the field
The day the dashboard went entirely empty. Someone changed stat_prefix to a more readable name. It had no effect on traffic, so the deployment passed quietly, and a few days later someone asks "was this graph always like this?" The values did not fall to 0; the time series itself disappeared, so an alert that does not treat "no data" as an outage never even fired. The name is an interface — before you change it, first find who uses it.
When there are too many time series. In a proxy with hundreds of clusters, the number of statistics runs to tens of thousands. The storage cost on the Prometheus side screams first, and the place to touch then is not the scrape interval but the list you export.
Reading a histogram like a counter. A value like request time almost always looks fine if you look at a single average. Slow requests are few in number and hardly move the average. That is why the histogram exists, and what you should look at is not the average but the upper percentiles. The text output of the admin port shows percentiles together, but if there are no recorded values it shows them as such, so you have to read "no value" and "zero" as different things.
"I reset it but the value did not change." That is mistaking a gauge for a counter. If you have the habit of telling the two kinds apart, this question never comes up.
Official documentation: Statistics overview · Administration interface
What you will do in the next lab
You confirm that the same request is counted under each of the three branches of names, pull out only what you need with filter and format=json, and see how names are split in the Prometheus output. Next you do these in turn: check that /reset_counters affects only counters, add a tag rule to create a label, and trim statistics with an exclusion list, and at the end you change the one word stat_prefix and see with your own eyes the statistics under the old name disappear entirely.