TT Lab
Get started
Learn Learning paths Courses

Istio Field Lab

Counted Twice, Paid Per Label

Continue in TT Lab

Goal

On a real mesh with the Prometheus addon up, ask for the standard metrics with PromQL, add and remove labels, turn off a metric, and filter the access log with the Telemetry API, and confirm the effects in the sidecar's metrics and logs.

Why it matters

Mesh observability looks free, but one label or one histogram multiplies the number of time series. Conversely, the dimensions the business needs are not in the standard labels. The Telemetry API is the device that decides what to add and what to throw away, and its effect has to be confirmed not from the configuration but from the metrics and logs that actually accumulate.

Steps

  1. Put up the materials with kubectl apply -f /opt/fixtures/istlab/telemetry-app.yaml and wait until they are ready. After calling http://web/ 10 times and http://web/broken 3 times from the client, ask Prometheus (Service istio-system/prometheus, port 9090) for sum by (response_code) (istio_requests_total{reporter="destination",destination_workload="web",destination_workload_namespace="shop"}) and save that response JSON as it is in /root/istlab-telemetry/01-prom.json. The scrape period is 15 seconds, so you wait until both codes are visible and then save.
  2. In Prometheus, sum istio_requests_total with destination_workload="web" and response_code="503" by reporter and write two lines into /root/istlab-telemetry/02-reporters.txt as source_503= and destination_503=.
  3. Write the Telemetry shop-telemetry of the namespace shop into /root/istlab-telemetry/telemetry.yaml and apply it — in metrics, the provider is prometheus, and the override adds tenant as request.headers['x-tenant'] through tagOverrides for REQUEST_COUNT and CLIENT_AND_SERVER. After applying, call three times with x-tenant: acme attached, and save the response to asking Prometheus for sum by (tenant) (istio_requests_total{tenant="acme"}) in /root/istlab-telemetry/03-tenant.json.
  4. Add request_protocol: {operation: REMOVE} to the tagOverrides of the same override in /root/istlab-telemetry/telemetry.yaml and apply again. After applying, the new time series of a request called with x-tenant: beta must not have the request_protocol label.
  5. Put one more override in /root/istlab-telemetry/telemetry.yaml, turn off REQUEST_DURATION (CLIENT_AND_SERVER) with disabled: true, and apply again. After applying, even if you send requests, the sum of the client sidecar's istio_request_duration_milliseconds_count must not increase (istio_requests_total must keep increasing).
  6. Add accessLogging to /root/istlab-telemetry/telemetry.yaml and apply again — provider envoy, filter.expression: "response.code >= 400". After applying, the client's 200 requests must not be left in the client sidecar access log, and only the /broken (503) request must be left.
  7. Write into /root/istlab-telemetry/07-error-ratio.promql a PromQL that returns the ratio of 5xx among the requests web received (reporter destination), and write the value that comes out of asking Prometheus that query into /root/istlab-telemetry/07-ratio.txt as ratio=.
  8. Write five lines into /root/istlab-telemetry/08-report.md — requests_503_destination= (step 2's destination_503), counted_twice= (yes if source and destination counted the same number in step 2), tenant_label= (the tenant value attached in step 3), duration_counting= (yes if the duration metric is increasing now), and logged_only_errors= (yes if it is the result of step 6) — and write what you learned below that in at least four lines.

Notes

Send requests and ask Prometheus

Put up the materials with kubectl apply -f /opt/fixtures/istlab/telemetry-app.yaml and wait until they are ready. After calling http://web/ 10 times and http://web/broken 3 times from the client, ask Prometheus (Service istio-system/prometheus, port 9090) for sum by (response_code) (istio_requests_total{reporter="destination",destination_workload="web",destination_workload_namespace="shop"}) and save that response JSON as it is in /root/istlab-telemetry/01-prom.json. The scrape period is 15 seconds, so you wait until both codes are visible and then save.

The sidecar raises standard metrics such as istio_requests_total on every request, and Prometheus looks at the Pod's annotations and scrapes the sidecar's port 15020. So it is not visible the instant you send a request but about one period later. You query through the HTTP API, as in curl -s http://<ClusterIP>:9090/api/v1/query --data-urlencode 'query=…'.

Who counts a single request — reporter

In Prometheus, sum istio_requests_total with destination_workload="web" and response_code="503" by reporter and write two lines into /root/istlab-telemetry/02-reporters.txt as source_503= and destination_503=.

In sidecar mode, a single request is counted separately by the sending-side (source) sidecar and the receiving-side (destination) sidecar. It is normal for the two values to come out equal, and if you add them on a dashboard without filtering the reporter, the request count looks doubled. Also remember that if the receiving side is outside the mesh, only the source side remains, and if the sending side is outside the mesh, only the destination side remains.

Attach our label to the metric — tagOverrides

Write the Telemetry shop-telemetry of the namespace shop into /root/istlab-telemetry/telemetry.yaml and apply it — in metrics, the provider is prometheus, and the override adds tenant as request.headers['x-tenant'] through tagOverrides for REQUEST_COUNT and CLIENT_AND_SERVER. After applying, call three times with x-tenant: acme attached, and save the response to asking Prometheus for sum by (tenant) (istio_requests_total{tenant="acme"}) in /root/istlab-telemetry/03-tenant.json.

The Telemetry API adjusts the sidecar's metric configuration per namespace (or workload). The value of tagOverrides is a request attribute expression, so you can make request information such as headers and paths into labels. But be careful: if you make a label out of values that grow endlessly (something like a user id), the time series explode. You can see what the sidecar attached right away with kubectl -n shop exec client -c istio-proxy -- pilot-agent request GET stats/prometheus, and it shows up in Prometheus about one period later.

Remove a label you do not use

Add request_protocol: {operation: REMOVE} to the tagOverrides of the same override in /root/istlab-telemetry/telemetry.yaml and apply again. After applying, the new time series of a request called with x-tenant: beta must not have the request_protocol label.

Even if a label always has the same value (http), it is attached to every time series and takes up storage. Removing a standard label you do not use is the cheapest way to reduce cardinality. Time series that have already been made are counters and are not erased but remain, so confirm with newly created time series (a new tenant value).

Turn off a metric you do not use

Put one more override in /root/istlab-telemetry/telemetry.yaml, turn off REQUEST_DURATION (CLIENT_AND_SERVER) with disabled: true, and apply again. After applying, even if you send requests, the sum of the client sidecar's istio_request_duration_milliseconds_count must not increase (istio_requests_total must keep increasing).

A latency histogram creates a time series for each bucket and is the heaviest of the standard metrics. With the Telemetry API you can make a choice such as looking at it only on the gateway and turning it off on sidecars. Turning it off does not make the old values disappear, so you confirm with 'it does not increase' — measure the sum twice, before and after the requests, and compare.

Leave only errors in the access log

Add accessLogging to /root/istlab-telemetry/telemetry.yaml and apply again — provider envoy, filter.expression: "response.code >= 400". After applying, the client's 200 requests must not be left in the client sidecar access log, and only the /broken (503) request must be left.

This mesh turned on all access logs through meshConfig at install time. The Telemetry accessLogging can override that per namespace so that only requests matching the condition expression (CEL) are left. If you leave even the successful requests, the log cost grows in proportion to the request count, so in production it is common to leave only errors and slow requests. Send with a request id attached and find it in kubectl -n shop logs client -c istio-proxy.

The error rate in one line of PromQL

Write into /root/istlab-telemetry/07-error-ratio.promql a PromQL that returns the ratio of 5xx among the requests web received (reporter destination), and write the value that comes out of asking Prometheus that query into /root/istlab-telemetry/07-ratio.txt as ratio=.

The ratio is sum(5xx 카운터) / sum(전체 카운터) (the terms are the 5xx counter and the total counter). If you do not pin the reporter to one, you divide counts where one request was counted twice, and if there are requests from outside the mesh, the denominator and numerator count different things. On a production dashboard you would look at the recent ratio with rate(…[5m]), but the requests in this lab are only a few, so you divide the counters themselves. You can send the line breaks of a query file as they are with --data-urlencode "query=$(cat 파일)" (the placeholder is the query file).

Write down what you counted and what you threw away

Write five lines into /root/istlab-telemetry/08-report.md — requests_503_destination= (step 2's destination_503), counted_twice= (yes if source and destination counted the same number in step 2), tenant_label= (the tenant value attached in step 3), duration_counting= (yes if the duration metric is increasing now), and logged_only_errors= (yes if it is the result of step 6) — and write what you learned below that in at least four lines.

In the explanation lines, write in your own words 'the cost of adding a label (cardinality)' and 'why a total that does not filter the reporter is wrong'.