Pull Just What You Need From Five Hundred Numbers
Goal
Divide the structure of statistic names into three branches, learn in turn how to pull them out, the rules for splitting them and how to trim them, and then confirm for yourself what disappears when you change a name.
Why it matters
Statistics are the floor that dashboards and alerts stand on. But that floor is made of names, so if you do not know the naming rules, you cannot carry a value you saw in /stats over to a dashboard query, and if you change a name casually, you empty whole graphs. On top of that, if you cannot tell counters from gauges, you get stuck at "I reset it but it did not change", and if you do not know the property of the exclusion list, you put it off with "I will turn it on when I need it" and find out, at the very moment you need it, that there are no past values.
Steps
- Start two upstreams —
8093isokand8094isfail. In/root/envd-stats/stats.yaml, put a listener whosestat_prefixisshopand clustersgoodandbad, and start it (admin 9971, listener127.0.0.1:10071). After you request/hellofive times and/bad/xonce, in/root/envd-stats/01-names.txtwrite three lines,cluster=,http=andlistener=— each is the full name of the statisticcluster.good.upstream_rq_total,http.shop.downstream_rq_total, and thedownstream_cx_totalstatistic of that listener. - Attach
filterto/statsto pull out only the statistics that containupstream_rq_total, addformat=jsonas well, and save it to/root/envd-stats/02-filter.json. Then in/root/envd-stats/02-filter.txt, write two lines:total_stats=(the number of lines received without a filter) andmatched=(the number of statistics narrowed down by the filter). - Fetch
/stats/prometheusand save only the lines corresponding tocluster.good.upstream_rq_totalto/root/envd-stats/03-prom.txt. Then in/root/envd-stats/03-prom.map, write three lines:envoy_name=(the metric name in Prometheus),label_key=(the key of the label that carries the cluster name) andvalue=(its value). - POST
/reset_counters, read the two statistics again, and in/root/envd-stats/04-reset.txtwrite four lines:counter_before=andcounter_after=(cluster.good.upstream_rq_total), andgauge_before=andgauge_after=(cluster.good.membership_total). - Create
/root/envd-stats/stats-tags.yaml— the same asstats.yaml, but withstats_config.stats_tagsat the top level, extracting the tag nameenvd_clusterwith the regular expression^cluster\.((.+?)\.). After you start it and request/hellotwice, save one line carrying theenvd_clusterlabel from/stats/prometheusto/root/envd-stats/05-tags.txt. - Create
/root/envd-stats/stats-trim.yaml— exclude statistics with the prefixcluster.bad.usingstats_config.stats_matcher.exclusion_list. After you start it, in/root/envd-stats/06-trim.txtwrite three lines:before=(the total number of statistics lines in the step 5 configuration),after=(the number of lines in this configuration) andbad_stats=(the number of lines starting withcluster.bad.in this configuration). - Create
/root/envd-stats/stats-rename.yaml— the step 1 configuration with onlystat_prefixchanged fromshoptocheckout. After you start it and request/hellotwice, in/root/envd-stats/07-rename.txtwrite three lines:old_prefix_stats=(the number of statistics lines starting withhttp.shop.),new_prefix_stats=(the number of lines starting withhttp.checkout.) andnew_rq_total=(the value ofhttp.checkout.downstream_rq_total). - In
/root/envd-stats/08-report.md, write five lines —counter_after_reset=andgauge_after_reset=(step 4),prom_label=(the tag name you added in step 5),trim_removed=(the before from step 6 minus the after) andrenamed_lost=(yes if the old-prefix statistics disappeared in step 7) — and below them write what you learned in at least four lines.
Notes
- If an admin port address contains
?or&, wrap the whole address in quotes. If you do not, the shell chops off what comes after. - Before you start Envoy again, clean up with
pkill -x envoy, and wait for startup with a loop that runs until/readyreturns LIVE. - Admin commands such as
/reset_countersand/quitquitquitare POST. Usecurl -X POST. - Statistic values are in the format
이름: 값(the placeholders are the name and the value), so pull them out withgrep -m1 '^이름:' | awk '{print $2}'(the placeholder is the statistic name). - Common mistake — writing the field name as
stat_tags. The correct name isstats_tags, and if you get it wrong, the whole configuration is rejected with "there is no such field".
Names start in three branches
Start two upstreams — 8093 is ok and 8094 is fail. In /root/envd-stats/stats.yaml, put a listener whose stat_prefix is shop and clusters good and bad, and start it (admin 9971, listener 127.0.0.1:10071). After you request /hello five times and /bad/x once, in /root/envd-stats/01-names.txt write three lines, cluster=, http= and listener= — each is the full name of the statistic cluster.good.upstream_rq_total, http.shop.downstream_rq_total, and the downstream_cx_total statistic of that listener.
Envoy statistic names state what the number is about with a prefix. cluster. is the upstream side, http. is the requests handled by that HTTP connection manager, and listener. is the socket side. One and the same request is counted in three places separately, so when the three numbers disagree, that difference becomes the clue. The listener name contains the address and port, and there is a spot where an underscore is used instead of a dot, so check for yourself.
Pull out only what you need from five hundred
Attach filter to /stats to pull out only the statistics that contain upstream_rq_total, add format=json as well, and save it to /root/envd-stats/02-filter.json. Then in /root/envd-stats/02-filter.txt, write two lines: total_stats= (the number of lines received without a filter) and matched= (the number of statistics narrowed down by the filter).
There are more than five hundred statistics even in the default configuration. That is why the admin port has query parameters — ?filter= is a regular expression and ?format=json converts it to a machine-readable shape. The two can be used together with &. In the shell, ? and & are special characters, so wrap the whole address in quotes. Count the JSON side with jq '.stats | length'.
The same value goes out under a different name
Fetch /stats/prometheus and save only the lines corresponding to cluster.good.upstream_rq_total to /root/envd-stats/03-prom.txt. Then in /root/envd-stats/03-prom.map, write three lines: envoy_name= (the metric name in Prometheus), label_key= (the key of the label that carries the cluster name) and value= (its value).
Prometheus has no concept of a "long name joined with dots". It expresses things as a metric name and a set of labels. So when Envoy exports, it splits the name — cluster.good.upstream_rq_total is divided into one metric name and a label carrying the cluster name. You need to know this rule to carry a name you saw in /stats over to a dashboard query. The labels are inside the braces as 열쇠="값" (the placeholders are the key and the value).
The reset button works on only half
POST /reset_counters, read the two statistics again, and in /root/envd-stats/04-reset.txt write four lines: counter_before= and counter_after= (cluster.good.upstream_rq_total), and gauge_before= and gauge_after= (cluster.good.membership_total).
Statistics come in kinds. A counter is a cumulative value that only keeps growing, a gauge is the state at this moment, and a histogram is the distribution of values. /reset_counters, as its name says, resets only counters to 0 — a gauge is a current state, such as "how many servers are healthy right now", so there is nothing to reset. If you do not know this difference, you get stuck at "I reset it but the value did not change". Send a POST with curl -X POST.
Decide for yourself the rules that split names
Create /root/envd-stats/stats-tags.yaml — the same as stats.yaml, but with stats_config.stats_tags at the top level, extracting the tag name envd_cluster with the regular expression ^cluster\.((.+?)\.). After you start it and request /hello twice, save one line carrying the envd_cluster label from /stats/prometheus to /root/envd-stats/05-tags.txt.
Envoy already holds several default tag rules (cluster name, listener address and so on). If you add a rule to them, you can pull your own organization's naming rules out as labels — for example, if you put a team name as a prefix on cluster names, you can make just that into a label, and then group by team on the dashboard. The first parenthesis of the regular expression is the part to cut out of the name, and the second parenthesis is the label value. The field name is stats_tags (not stat_tags).
Choose what not to export
Create /root/envd-stats/stats-trim.yaml — exclude statistics with the prefix cluster.bad. using stats_config.stats_matcher.exclusion_list. After you start it, in /root/envd-stats/06-trim.txt write three lines: before= (the total number of statistics lines in the step 5 configuration), after= (the number of lines in this configuration) and bad_stats= (the number of lines starting with cluster.bad. in this configuration).
Statistics eat memory, and when you export to Prometheus, the storage cost is as many as the number of time series. In a proxy with hundreds of clusters, this number quickly becomes hard to bear. So there is a mechanism that chooses what to export — you specify prefixes, exact names or regular expressions with an exclusion list or an inclusion list. The thing to watch out for is that an excluded statistic does not disappear from the display; it is not recorded at all. Even if you bring it back after you need it, there are no values from the time in between.
Change one word and the dashboard goes empty
Create /root/envd-stats/stats-rename.yaml — the step 1 configuration with only stat_prefix changed from shop to checkout. After you start it and request /hello twice, in /root/envd-stats/07-rename.txt write three lines: old_prefix_stats= (the number of statistics lines starting with http.shop.), new_prefix_stats= (the number of lines starting with http.checkout.) and new_rq_total= (the value of http.checkout.downstream_rq_total).
stat_prefix is a value for names, so it has no effect on traffic. That makes it easy to change casually while refactoring, but the moment you do, every dashboard and alert that used that prefix becomes an empty graph. What is more, the values do not fall to 0; the time series itself disappears, so if an alert does not treat "no data" as an outage, nobody notices. The name is an interface — when you change it, you have to find who uses it first.
Leave a statistics operations memo
In /root/envd-stats/08-report.md, write five lines — counter_after_reset= and gauge_after_reset= (step 4), prom_label= (the tag name you added in step 5), trim_removed= (the before from step 6 minus the after) and renamed_lost= (yes if the old-prefix statistics disappeared in step 7) — and below them write what you learned in at least four lines.
This memo is something you will read yourself the next time you build a dashboard or refactor a proxy configuration. It is good to write the last line in particular as the rule "the name is an interface" — take the values from the files you made in the earlier steps.