TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

What you gain and what you must split when the proxy leaves the pod

Continue in TT Lab

In one line

Ambient moves the proxy that used to sit in every Pod to a per-node L4 proxy, and leaves only the features that need to see the request contents to a separate L7 proxy called a waypoint — so migration starts not from one line of label but from a list of features.

Why another data plane was needed

The sidecar model has three costs. First, one more container per Pod. With 2,000 Pods that is 2,000 proxies and that much reserved request resource. Second, to add, remove, or upgrade the proxy, you must restart the Pods. Joining the mesh becomes a deployment event, so adoption itself is tied to each team's deployment schedule. Third, because the proxy is in the same Pod as the application, problems such as startup order, shutdown order, and resource contention come along.

Ambient avoids all three at once. It lays down two things on each node — a component that intercepts traffic from the Pod, and an L4 proxy that carries the intercepted traffic over a mutually authenticated tunnel. You leave the Pods alone, and with one line of namespace label they join or leave the mesh. No restart is needed.

Features split into two layers

There is something you must know exactly here. The node's L4 proxy does not look at the contents of a request. So what works and what does not split clearly.

Feature With the L4 proxy alone Needs a waypoint
Mutual authentication and encryption between workloads Works —
Allow/deny by caller identity (L4) Works —
Policy by port and identity Works —
HTTP path- and header-based routing Does not work Needed
Authorization per method and path Does not work Needed
Retries, timeouts, and traffic splitting Does not work Needed

A waypoint is expressed not as an Istio-specific object but as a Gateway of the Gateway API. A gateway class of istio-waypoint is what makes it a waypoint, and its listener accepts as it is the tunnel protocol that ambient uses between Pods. What the waypoint is for is told by the label istio.io/waypoint-for — whether to put it in front of services, in front of workloads, in front of both, or in front of nothing.

How the switch is done

The trap is that there are two labels. The sidecar uses the label that the injection webhook looks at, and ambient uses the label that the node's component looks at. The two do not know about each other. So even if you put both on a namespace, no error occurs, and it becomes a state in which it is undecided which data plane the Pods of that namespace use. istioctl analyze catches this case as IST0123 — which is why you put this command into the checklist of the switchover work.

Set the order like this. First write down the mesh features that namespace actually uses. If there is even one L7 feature, put up the waypoint first. Then remove the sidecar label and apply the ambient label. Reverting is the reverse order. During the switchover the two modes coexist in one cluster, so it is better to build a script that counts which namespace is in which mode than to rely on human memory.

What it looks like in the field

The most common incident is "I moved to ambient and header-based routing does not work." The rule is still there and there is no error. It is just that the L7 proxy that would enforce that rule is not in that place. If you move without knowing the two layers, you will certainly run into it.

What makes this incident especially nasty is that it is quiet. The L4 layer keeps working, so connections are made and mutual authentication applies. The only thing cut is the paths that used to split by request content, so on the surface the mesh runs well but only the traffic that should go to a particular version goes somewhere else. No error is left in the logs either. That is why the switchover plan must include a "list of the mesh objects this namespace uses," and if even one of them requires L7, you put up the waypoint first and then change the label.

The second is when you add a label without removing the old one. If the sidecar label remains, a proxy is still injected into newly created Pods. On screen it looks as if the switch is complete, but in fact it is only half moved, and this state changes a little with each deployment.

Limits of this lab environment

In this environment you cannot actually bring up either ztunnel or a waypoint. So you cannot see traffic passing through the tunnel or L4 authorization actually being enforced. On top of that, this cluster has no Gateway API CRD, so a waypoint manifest can be generated but cannot be applied. We confirm this fact directly within the lab instead of hiding it — because the error message of the failed application says most clearly what a waypoint stands on. Instead, profile rendering, labels and their conflict detection, and the fields of the waypoint manifest can all be confirmed with real tools and a real API server.

What you will do in the next lab

You render the ambient profile and compare its components and identities with the sidecar profile, and see what more is turned on in the mesh configuration. You create the two modes with namespace labels and check what the analyzer catches when you layer them, and then generate a waypoint manifest, read its fields, and even see why the application fails. At the end you build a checklist that counts the cluster's namespaces by mode.