TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

Metrics, Access Logs and Traces — What the Proxy Gives You Free, and What It Does Not

Continue in TT Lab

In one line

While a request passes through the proxy anyway, the mesh produces metrics, access logs, and spans for you. Only traces, however, break unless the application passes the headers along.

Why this was needed

If you handle observability in the application, two things go out of step. First, the definitions differ per service. Some teams count only 5xx as errors and some also count 4xx. Some teams measure latency from the client's point of view and some from the server's. Even if you put the dashboards side by side, they cannot be compared. Second, services that are not instrumented remain. Internal tools made in a hurry, the legacy of a team you took over, and components received from an outsourcer have no instrumentation, and those are always the blind spots of outages.

A mesh solves this problem with "the request passes through the proxy anyway, so let us measure there." The names and labels of the standard metrics become the same in every service. This uniformity is of far more value than attaching instrumentation anew.

How it works

Metrics. Each sidecar exposes statistics in Prometheus format on port 15090. The key ones are three.

Metric Type Where it is used
istio_requests_total Counter Request count, error rate
istio_request_duration_milliseconds Histogram p50/p99 latency
istio_request_bytes / istio_response_bytes Histogram Payload size

Among the labels there are three worth knowing in particular. reporter is split into source (the sending side's proxy) and destination (the receiving side's proxy), and the same request is recorded on each side, so if you add the two together you count double. response_flags has the same values as the flags in the access log, so you can tell the nature of a 503 from metrics alone. connection_security_policy is mutual_tls or none, and the completion criterion for moving to STRICT is exactly zero requests with this label at none.

One trap is cardinality. If you put values such as the request path or user ID in a label, the time series explode. For custom tags, put in only values with a finite set of kinds.

It is also worth knowing the history of how the architecture changed. In the past, a separate service called Mixer was called on every request to collect metrics, and that was a cause of latency and outages. Now filters inside the proxy produce them directly (Telemetry API v2). So observability no longer adds a hop to the request path.

Access logs. The proxy leaves one line per request. The parts a person needs to read are the response code and the flag next to it.

Flag Meaning What to look at first
UH No healthy upstream Subset labels and Pod labels
UF Upstream connection failure An mTLS mismatch where only one side is STRICT
UO Connection pool overflow The circuit breaker limit
NR No route Whether a catch-all route exists
URX Retries exhausted The root cause is with the other flags
UAEX External authorization denied The CUSTOM policy and the state of the authorizer

Logging everything is expensive. A common setup is to put a condition such as response.code >= 400 in a Telemetry resource's filter and keep only failures.

Distributed tracing. This is the "not free" part. The proxy even creates spans, measures timing, and sends them to the collector. But copying the inbound request's trace headers (the B3 family or W3C traceparent) onto outbound requests is something the application must do. If it does not, a separate trace is created for each service and the call chain breaks. The answer to "I installed the mesh but traces only look one hop long" is almost always this.

The default sampling rate is 1%. The first proxy decides whether to sample and propagates it in headers, so the later ones follow that decision. You raise it only when debugging and do not keep 100% permanently.

Visualization. Kiali draws a service graph from Prometheus metrics, and reads the Kubernetes API and Istio configuration and even checks referential integrity. That is, Kiali's configuration validation shows on screen the same kind of work as istioctl analyze.

What it looks like in the field

First, a dashboard that counts source and destination mixed together. If the request count comes out exactly double, nine times out of ten you left out the reporter filter. You must fix it to one of source for a client view or destination for a server view.

Second, mesh metrics cannot replace app metrics. The proxy knows only that "the request returned 200." It does not know whether the response body was wrong or whether the order was actually saved. Mesh observability is a uniform floor at the infrastructure layer, and domain metrics must still be produced by the application.

Third, paths without a sidecar are invisible. Traffic from crons outside the mesh or from namespaces without injection does not appear in the graph at all. A clean graph does not mean there is no traffic.

What you will do in the next lab

Let us say this honestly first. This lab environment has no Envoy that actually carries traffic. Even when a Pod with a sidecar becomes Running, requests do not flow, and you can neither scrape 15090 nor hit the admin API on 15000. So we did not build a lab that fabricates metric values or trace pictures.

Instead, the lab right after covers the side that designs signals. You lay down a default access log for the whole mesh, apply a filter that keeps only failures by response code, attach a sampling rate and literal tags to one workload, and turn off one standard metric. Then you build a checker that filters out, before deployment, tags that blow up cardinality, and you see that istioctl analyze catches it as an error if there are two Telemetry resources without selectors in one namespace. Even when no values flow, whether the configuration is valid and whether the scope is right are verified as they are, and the places where people get caught in the field are mostly those two.