TT Lab
Get started
Learn Learning paths Courses

ICA — Istio Certified Associate

What Changes When You Detach the Proxy From the Pod

Continue in TT Lab

In one line

Ambient mode is a structure that separates L4 from L7 so that not every workload is forced to pay the full Envoy cost. In exchange, the failure domain grows from the Pod to the node.

Why this was needed

The cost of the sidecar model is clear in numbers. In a cluster of 500 Pods, if each sidecar uses 120MiB of memory, that alone is about 58GiB, and with a CPU request of 100m each, 50 vCPUs vanish from scheduling capacity. On top of that, since the proxy lives inside the Pod, raising the version requires a rolling restart of every workload, and if just one of the injection label, revision tag, or webhook configuration is off, you end up in a state where "some Pods are in the mesh and some are outside it." If you turn on STRICT in this state, every Pod without injection is cut off.

One more decisive observation is added to this. The services that actually need L7 features are a fraction of the whole. For the vast majority of the rest, mTLS and basic telemetry are enough, yet the sidecar model imposed a full Envoy on them too.

How it works

Ambient is split into two layers.

Traffic between nodes flows through an HBONE tunnel. It wraps the original TCP stream in HTTP/2 CONNECT over mTLS on port 15008. Thanks to HTTP/2 stream multiplexing, several connections between the same pair of nodes share one mTLS connection, reducing handshake cost, and since the original destination and identity are carried in the CONNECT header, the receiving ztunnel can decide L4 authorization without opening the payload.

You must remember one design principle. A waypoint is owned by the destination. In the sidecar model, the client-side proxy did routing and retries, but in ambient, the team that owns the destination service enforces L7 policy at its own waypoint. Policy ownership becomes clear, but in exchange client-side policies such as "a different timeout per caller" need to be redesigned.

You must honestly look at the trade-off too. A path through a waypoint has 3 hops (ztunnel → waypoint → ztunnel), so it can be longer than a sidecar (2 hops). And if a ztunnel dies, all mesh traffic on that node is affected. With sidecars, a proxy failure was confined to one Pod, but ambient has a node-level failure domain. You must reflect a PodDisruptionBudget, a priority class, and restart monitoring in your operational design.

In debugging, order is everything. You read istioctl proxy-config in the order the traffic flows.

listener  이 프록시가 그 포트를 듣고 있는가
   ->
route     VirtualService 가 라우팅 테이블로 번역됐는가
   ->
cluster   DestinationRule(서킷브레이커, TLS)이 반영됐는가
   ->
endpoint  그 subset 에 실제 파드 IP 가 잡혀 있는가

If you find at which step things diverge from expectation, the resource to fix is automatically determined. The table for reading 503 response flags is in the same vein.

Flag Meaning First suspect
UH no healthy upstream Subset label and Pod label mismatch
UO upstream overflow connectionPool limit exceeded
UF upstream connection failure An mTLS mismatch where only one side is STRICT, or a port protocol misidentified
NR no route Missing VirtualService match, no catch-all
URX max retries reached Retries exhausted — trace the root cause together with other flags

One last thing on observability. Envoy creates spans automatically, but if the application does not propagate the inbound request's trace headers to outbound, the trace is cut. This is the reason traces come out in pieces after you install a mesh, and it is also the only part the mesh cannot do for you.

What it looks like in the field

In the author's homelab, hubble-relay and hubble-ui once stayed Pending. The event was 0/1 nodes are available: 1 node(s) had untolerated taint(s). The cause was simple. The control plane node has the taint node-role.kubernetes.io/control-plane:NoSchedule, and these two components are Deployments, not DaemonSets, and so had no toleration. This contrasted with CoreDNS, which has a default toleration and started normally, and it was resolved right away when a worker joined.

This case is directly useful for understanding ambient. ztunnel is a DaemonSet and a waypoint is a Deployment. That is, a ztunnel comes up automatically on every node, but a waypoint needs the scheduler to find it a place. If only tainted nodes remain or resources are insufficient, a waypoint silently stays Pending, and the L7 policies of that namespace are not enforced at all. This is why the first place to check for the symptom "L4 works but only L7 policy doesn't take effect" is the waypoint Pod's status.

One more thing. That homelab was built with --skip-phases=addon/kube-proxy so that kube-proxy was never installed in the first place, rather than installed and then deleted. This is because once it has run, the KUBE-SERVICES / KUBE-SVC-* / KUBE-SEP-* chains remain on the node. The same principle applies when moving from sidecar to ambient. The injection label and the ambient label must not coexist on one namespace, and after removing the injection label you must be sure to clear out the existing sidecars with a rollout restart before attaching the ambient label.

What to check in the next quiz

This module is a concepts module. Instead of a lab, the quiz checks the ambient structure, HBONE, failure domains, the order of reading proxy-config, and reading response flags. These five are the tools you reach for first not only on the exam but in real incident response.