TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

Design Your Signals with the Telemetry API

Continue in TT Lab

Goal

You learn how to shape by configuration the signals the proxy produces. You lay down a default log for the whole mesh, apply a condition that keeps only failures, attach sampling and tags to one workload, and go as far as filtering out, before deployment, tags that blow up cardinality.

Why it matters

What a mesh sells is uniformity. Even a service with no instrumentation gets metrics with the same names and the same labels, and that uniformity makes dashboards comparable. But if you leave the defaults as they are, two things become problems. An access log is one line per request, so storage cost quickly grows for a service with a lot of traffic, and if you put in a custom tag wrongly, the time series grow by the number of distinct label values and storage and queries collapse together.

The Telemetry API lets you make that adjustment declaratively. The key is that there are three layers of scope. If you put it in the installation namespace without a selector, it is the whole mesh; if you put it in an ordinary namespace without a selector, it is that namespace; and if you attach a selector, it is that workload. And there must be only one without a selector per namespace. If there are two or more, it is undefined which one wins.

Environment

This Pod has neither a real istiod nor a sidecar. Requests do not flow, so neither metrics nor logs are produced. Instead there is a real apiserver with the Telemetry CRD registered and istioctl, so whether the configuration is valid and whether the scope is right are verified as they are. Those two are also the places where people get caught in the field. The working directory is /root/istio-obs, and you use k8s/, reject/, bin/, and out/ under it.

Steps

  1. Create mesh-obs and lay down the mesh default access log in istio-system.
  2. In mesh-obs, apply a log filter that keeps only failures.
  3. Apply 10% sampling and a literal tag to app=reviews, and see out-of-range values get blocked.
  4. Apply metric tuning to app=ratings.
  5. Filter out dangerous tags before deployment with bin/tag-lint.sh.
  6. Create two workloads and check the scope rule with istioctl analyze.
  7. In out/effective.txt, calculate and write what actually applies to one Pod.
  8. Summarize it in six lines in out/summary.txt.

Notes

Lay down the default access log for the whole mesh

Create the mesh-obs namespace and attach the label istio-injection=enabled, and create the Telemetry mesh-default in istio-system as /root/istio-obs/k8s/mesh-default.yaml. Specify only the accessLogging provider, without a selector.

A Telemetry's scope has three layers. In istio-system (the installation namespace) without a selector it is the whole mesh, in an ordinary namespace without a selector it is that namespace, and with a selector it is that workload. The fact that having no selector itself decides the scope is the key of this API. The provider name is a string that points to a name registered in the mesh configuration, so here the format is valid even if it does not actually exist.

Keep only failed requests in the log

Create the Telemetry failures-only in mesh-obs as /root/istio-obs/k8s/failures-only.yaml. Do not put a selector, and set on accessLogging a filter.expression that keeps only when the response code is 400 or higher.

Logging everything is expensive. The proxy leaves one line per request, so storage cost quickly grows for a service with a lot of traffic. The filter is a CEL expression, and in the access log context you can use response.code. This Telemetry applies to the whole namespace, so you must not attach a selector. And there must be only one Telemetry without a selector per namespace, so always attach selectors to the ones in later steps.

Apply a sampling rate and custom tags to a workload

Create the Telemetry trace-sampling in mesh-obs as /root/istio-obs/k8s/trace-sampling.yaml. selector.matchLabels.app is reviews, tracing[0].randomSamplingPercentage is 10, and put one tag with a literal value in customTags. Then write an out-of-range sampling value in /root/istio-obs/reject/bad-sampling.yaml, try to apply it, and save the rejection message to /root/istio-obs/out/rejected.txt.

The default sampling rate is 1%. The first proxy decides whether to sample and propagates it in headers, so later hops follow that decision. You raise it only when debugging and do not keep 100% permanently. For custom tags you must put in only values with a finite set of kinds, so use a literal here. An out-of-range value is blocked by the CRD schema, so the object is not created at all. The message goes out on standard error, so capture it with 2>&1.

Turn off a standard metric and add a tag

Create the Telemetry metrics-tuning in mesh-obs as /root/istio-obs/k8s/metrics-tuning.yaml. selector.matchLabels.app is ratings, and the provider is prometheus. With overrides, turn off REQUEST_SIZE by specifying a mode, and in another item add one more tag with a literal value using tagOverrides.

Turning a metric off is also configuration. If you turn off a histogram you do not use, the time series shrink and both storage and queries get lighter. mode is chosen from CLIENT, SERVER, and CLIENT_AND_SERVER, and because the same request is recorded by the proxy on each side, you must decide which side you mean. A tag value is a CEL expression, so to put in a constant string you must wrap it in single quotes.

Filter out tags that blow up cardinality before deployment

Create /root/istio-obs/bin/tag-lint.sh <Telemetry파일> (the placeholder is the Telemetry file). If the value of a tagOverrides entry points to something with a nearly unbounded set of kinds of values, such as the request path, a request identifier, a user identifier, or a request header, print HIGH_CARDINALITY=<태그이름>=<값> (the placeholders are the tag name and the value) and end with exit code 1. For a file that uses only constant strings, print OK and end with 0.

The number of time series grows as the product of label value combinations. If you use something with a nearly unbounded set of values, such as a user identifier, as a tag, storage and queries collapse together, and there is no safeguard in which the proxy hashes it or lowers sampling for you. If the CEL value is a constant wrapped in single quotes it is safe, and if it points to something like request.url_path or request.headers[...] it is dangerous. The grader actually runs this checker against a safe file and a dangerous file, and also checks that the file you made in step 4 does not get caught.

Check the scope rule with static analysis

Create the Deployments reviews and ratings in mesh-obs as /root/istio-obs/k8s/workloads.yaml. The app Pod label of each must be the same as its name. Then check that istioctl analyze -n mesh-obs ends with no errors.

A selector looks at Pod labels, not the Deployment name. So you must match the labels of the Pod template exactly. And if there are two or more Telemetry without a selector in one namespace, it is undefined which one wins, and istioctl analyze catches it as an error. If you left out a selector in an earlier step, it shows up here. This Pod has no real sidecar, so warnings related to injection remain, but if there are no errors it passes.

Calculate what actually applies to one Pod

Calculate which Telemetry applies to the app=reviews Pod and write it in five lines in /root/istio-obs/out/effective.txt. They are MESH, NAMESPACE, WORKLOAD, TRACING_PERCENT, and METRICS_APPLIES.

The three layers apply overlapping. One at the mesh scope, one at the namespace scope, and one at the workload scope that selects that Pod. For the first three, write the name, and for TRACING_PERCENT, write the sampling value written in the workload scope. For the last line, write the answer to whether the metric configuration from step 4 applies to this Pod too, as yes or no. Just look again at which app the selector was selecting. Do not guess the names; check and write them with kubectl get telemetry -A.

Count what you made and check that everything is valid

Check that all the files in /root/istio-obs/k8s pass istioctl validate, and summarize in six lines in /root/istio-obs/out/summary.txt. They are TELEMETRY_TOTAL, MESH_SCOPED, NAMESPACE_SCOPED, WORKLOAD_SCOPED, SAMPLING_PERCENT, and TRACE_HEADER_PROPAGATION.

For the first five lines, count from the cluster and write them. The last line is the answer to the fact that most often catches people in this course. The proxy even creates spans, measures timing, and sends them to the collector, but write in one word who must do the work of copying the inbound request's trace headers onto outbound requests. If that is not done, a separate trace is created for each service and the call chain breaks.