Metrics Are Counted Twice, and Labels Cost Money
In one line
The sidecar raises standard metrics (istio_requests_total, istio_request_duration_milliseconds, and so on) on every request and leaves an access log. The Telemetry API is a device that adjusts them per namespace or workload — it adds labels (tagOverrides), removes them (REMOVE), turns metrics off (disabled), and filters the log by a condition (filter).
Why this was needed
One of the mesh's promises is 'without changing code, you see the request count, error rate, and latency of every service in the same shape'. It is possible because the sidecar sits on every request. But if you leave it as it is, two problems arise.
First, cost. The standard metrics carry several labels such as source, destination, response code, and protocol, and the latency histogram creates a time series for each bucket. With hundreds of services, Prometheus gets overwhelmed first. Second, shortfall. A dimension the business needs, such as 'which customer (tenant) is this request from', is not in the standard labels.
The Telemetry API handles both with a single setting. You can apply it narrowing down in the order of the whole mesh (istio-system), a namespace, and a workload selector.
How it works
Metrics are counted twice. In sidecar mode, one request is counted by the sending sidecar (reporter="source") and the receiving sidecar (reporter="destination") each. If both are inside the mesh, the same count piles up twice, so when you take a total you have to pin the reporter to one. The receiving side (destination) is usually closer to 'the requests the service received'.
Prometheus scrapes. The distribution's Prometheus addon looks at the Pod annotations and periodically (this setup is 15 seconds) scrapes the sidecar's merged metrics port. So after you send a request, it shows up in a query result about one period later. You can see the value the sidecar holds right now directly with pilot-agent request GET stats/prometheus.
What you adjust with Telemetry
| Setting | What it does |
|---|---|
tagOverrides: {tenant: {value: "request.headers['x-tenant']"}} |
Makes a new label from a request attribute |
tagOverrides: {request_protocol: {operation: REMOVE}} |
Removes a standard label |
match: {metric: REQUEST_DURATION}, disabled: true |
Does not produce that metric |
accessLogging: [{filter: {expression: "response.code >= 400"}}] |
Leaves a log only for requests that match the condition |
If you make a label out of values that grow endlessly (a user id, a request id), the time series explode. When you add a label, first think about the number of distinct values. And counters are not erased, so the effect of a configuration change shows only in newly created time series.
You get a ratio with PromQL. The error rate is sum(5xx) / sum(전체) (the denominator is the total). A production dashboard looks at the ratio over a recent window with rate(…[5m]), but in an experiment with few requests you can just divide counter by counter. In both cases the numerator and denominator must point at the same reporter and the same destination.
What it looks like in the field
"The request count on the dashboard is twice the load balancer's." You added them without filtering the reporter.
"I added a tenant label and Prometheus memory jumped." There were tens of thousands of customers. It is a dimension you should look at in logs or traces, not in labels.
"I reduced the access log and failure analysis got harder." If you keep only errors, you cannot see 'since when did the requests that were normal get slow'. It is common to put the latency (response.duration) into the condition expression too.
Official docs: Telemetry API · Customizing Istio Metrics · Istio Standard Metrics · Envoy Access Logs · Prometheus addon
What you will do in the next lab
In a mesh that has the distribution's Prometheus addon installed with it, you send requests and ask about istio_requests_total through the HTTP API, and see the same request counted twice as source and destination. With the Telemetry API you add a tenant label, remove a label you do not use, turn off the latency metric, and filter the access log to keep only errors, and finally get the error rate in a single line of PromQL.