ICA — Istio Certified Associate
What Changes When You Detach the Proxy From the Pod
In one line
Ambient mode is a structure that separates L4 from L7 so that not every workload is forced to pay the full Envoy cost. In exchange, the failure domain grows from the Pod to the node.
Why this was needed
The cost of the sidecar model is clear in numbers. In a cluster of 500 Pods, if each sidecar uses 120MiB of memory, that alone is about 58GiB, and with a CPU request of 100m each, 50 vCPUs vanish from scheduling capacity. On top of that, since the proxy lives inside the Pod, raising the version requires a rolling restart of every workload, and if just one of the injection label, revision tag, or webhook configuration is off, you end up in a state where "some Pods are in the mesh and some are outside it." If you turn on STRICT in this state, every Pod without injection is cut off.
One more decisive observation is added to this. The services that actually need L7 features are a fraction of the whole. For the vast majority of the rest, mTLS and basic telemetry are enough, yet the sidecar model imposed a full Envoy on them too.
How it works
Ambient is split into two layers.
- ztunnel — a DaemonSet that runs one per node. It is a lightweight proxy written in Rust and is not Envoy. Its role is deliberately limited to L4. It covers mTLS, L4 authorization based on SPIFFE identity, and TCP telemetry. Since it does not parse HTTP, its memory scales with the number of connections, not the number of Pods.
- waypoint — an Envoy-based proxy deployed only per namespace or service account that needs L7. It is declared as a Gateway resource of the Gateway API, and as an ordinary Deployment it is scaled separately with an HPA.
Traffic between nodes flows through an HBONE tunnel. It wraps the original TCP stream in HTTP/2 CONNECT over mTLS on port 15008. Thanks to HTTP/2 stream multiplexing, several connections between the same pair of nodes share one mTLS connection, reducing handshake cost, and since the original destination and identity are carried in the CONNECT header, the receiving ztunnel can decide L4 authorization without opening the payload.
You must remember one design principle. A waypoint is owned by the destination. In the sidecar model, the client-side proxy did routing and retries, but in ambient, the team that owns the destination service enforces L7 policy at its own waypoint. Policy ownership becomes clear, but in exchange client-side policies such as "a different timeout per caller" need to be redesigned.
You must honestly look at the trade-off too. A path through a waypoint has 3 hops (ztunnel → waypoint → ztunnel), so it can be longer than a sidecar (2 hops). And if a ztunnel dies, all mesh traffic on that node is affected. With sidecars, a proxy failure was confined to one Pod, but ambient has a node-level failure domain. You must reflect a PodDisruptionBudget, a priority class, and restart monitoring in your operational design.
In debugging, order is everything. You read istioctl proxy-config in the order the traffic flows.
listener 이 프록시가 그 포트를 듣고 있는가
->
route VirtualService 가 라우팅 테이블로 번역됐는가
->
cluster DestinationRule(서킷브레이커, TLS)이 반영됐는가
->
endpoint 그 subset 에 실제 파드 IP 가 잡혀 있는가
If you find at which step things diverge from expectation, the resource to fix is automatically determined. The table for reading 503 response flags is in the same vein.
| Flag | Meaning | First suspect |
|---|---|---|
| UH | no healthy upstream | Subset label and Pod label mismatch |
| UO | upstream overflow | connectionPool limit exceeded |
| UF | upstream connection failure | An mTLS mismatch where only one side is STRICT, or a port protocol misidentified |
| NR | no route | Missing VirtualService match, no catch-all |
| URX | max retries reached | Retries exhausted — trace the root cause together with other flags |
One last thing on observability. Envoy creates spans automatically, but if the application does not propagate the inbound request's trace headers to outbound, the trace is cut. This is the reason traces come out in pieces after you install a mesh, and it is also the only part the mesh cannot do for you.
What it looks like in the field
In the author's homelab, hubble-relay and hubble-ui once stayed Pending. The event was 0/1 nodes are available: 1 node(s) had untolerated taint(s). The cause was simple. The control plane node has the taint node-role.kubernetes.io/control-plane:NoSchedule, and these two components are Deployments, not DaemonSets, and so had no toleration. This contrasted with CoreDNS, which has a default toleration and started normally, and it was resolved right away when a worker joined.
This case is directly useful for understanding ambient. ztunnel is a DaemonSet and a waypoint is a Deployment. That is, a ztunnel comes up automatically on every node, but a waypoint needs the scheduler to find it a place. If only tainted nodes remain or resources are insufficient, a waypoint silently stays Pending, and the L7 policies of that namespace are not enforced at all. This is why the first place to check for the symptom "L4 works but only L7 policy doesn't take effect" is the waypoint Pod's status.
One more thing. That homelab was built with --skip-phases=addon/kube-proxy so that kube-proxy was never installed in the first place, rather than installed and then deleted. This is because once it has run, the KUBE-SERVICES / KUBE-SVC-* / KUBE-SEP-* chains remain on the node. The same principle applies when moving from sidecar to ambient. The injection label and the ambient label must not coexist on one namespace, and after removing the injection label you must be sure to clear out the existing sidecars with a rollout restart before attaching the ambient label.
What to check in the next quiz
This module is a concepts module. Instead of a lab, the quiz checks the ambient structure, HBONE, failure domains, the order of reading proxy-config, and reading response flags. These five are the tools you reach for first not only on the exam but in real incident response.