TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

Sidecars — Changing the Network Without Touching the Code

Continue in TT Lab

In one line

A service mesh is not a new network; it is a structure that attaches one proxy to every Pod and configures those proxies centrally.

Why this was needed

Once you have built around ten microservices, the same code shows up repeatedly on every team: retries, timeouts, circuit breakers, loading TLS certificates, verifying the peer, request logging, and propagating trace headers. If services are scattered across Java, Go, Python, and Node, you have to implement this list four times, and the four implementations behave subtly differently. To change a single default timeout, you have to open PRs in four repositories and deploy four times.

The idea of a mesh is simple. Take those common concerns out of the application process and have a proxy in the same Pod handle them instead. The application still sends a plaintext request to http://reviews:9080, but what actually carries that request to the other side is the proxy. So turning on mTLS becomes a matter of applying one configuration resource rather than deploying applications.

How it works

Istio is split into two layers. The control plane (istiod) and the data plane (the Envoy in each Pod). istiod merged Pilot, Citadel, and Galley into a single binary in 1.5, and Mixer, which used to sit in the request path, was removed entirely in 1.8. That is, nothing asks the control plane while a single request is being processed. Even if istiod dies, an Envoy that has already received its configuration keeps carrying traffic with the last configuration.

Configuration flows over gRPC streams called xDS.

API What it carries Istio-side input
LDS Listeners (ports, protocols) Port names, Gateway
RDS HTTP routing rules VirtualService
CDS Upstream clusters DestinationRule
EDS The actual endpoints of a cluster Kubernetes Endpoints
SDS TLS certificates and keys The istiod CA

Istio sends all of these merged into one stream called ADS. This is to prevent the configuration from momentarily breaking because endpoints arrive before the cluster definition.

Injection is done by Kubernetes' MutatingAdmissionWebhook. If a namespace carries the label istio-injection=enabled, then at the moment a Pod is created in that namespace, istiod modifies the Pod spec and inserts two containers.

There is a spot here where incidents often happen. The label works only at Pod creation time. Nothing happens to Pods that are already running, so once you attach the label you must trigger a rollout to recreate the Pods before the sidecar gets in. The answer to "I attached the label, so why are they not in the mesh?" is this nine times out of ten.

What it looks like in the field

First, you see the price of a sidecar as a bill. Roughly 100MB of memory and a few seconds of startup time are added per Pod. With 100 Pods that is 10GB. That is why Ambient mode appeared, in which a per-node ztunnel handles L4 and a per-namespace waypoint handles L7 without sidecars, and configurations that go in an entirely different direction, choosing an eBPF datapath (like Cilium) that even removes kube-proxy, have become common. Adopting a mesh should mean that you can answer "what do we buy by paying this cost?"

Second, an upgrade is a restart. To raise the sidecar version you have to recreate every Pod. On a cluster with thousands of Pods this is a job of several hours, which is why a mesh upgrade plan must include a rollout order and a maintenance window.

Third, istioctl kube-inject is the offline edition of the webhook. It applies in advance to a manifest file what the webhook does at Pod creation and shows you the result. You use it when you want to pin the injection result with GitOps, and, as now, when you want to see the anatomy of the injection result with your own eyes.

What you will do in the next lab

You generate the installation manifest with istioctl manifest generate, count what is in it, and register the mesh API types (CRDs) in the cluster. You attach the injection label to mesh-lab and not to legacy, creating a control group. Then you run istioctl kube-inject on the fixture workload and summarize in JSON how many containers there are and what the init container does.

There is something to state up front. In this lab environment, a real Envoy does not carry traffic. Pods become Running but packets do not flow, and there is no kubectl exec or port forwarding either. So all the grading in this course looks at writing manifests and static verification with istioctl analyze/validate. Actual handshakes and metrics are covered by reading and quizzes. The eye for judging whether configuration is right is developed well enough this way, and in the field too, half of incidents come from configuration, not packets.