TT Lab
Get started
Learn Learning paths Courses

ICA — Istio Certified Associate

An Unenforced State Looks Perfectly Normal

Continue in TT Lab

In one line

The most dangerous thing in a mesh is not a wrong configuration but a configuration that has been applied nowhere. The former produces an error, but the latter looks normal.

Why this is a problem

The most dangerous state in a service mesh is not "the configuration is wrong" but "the configuration is applied nowhere." The former produces an error, but the latter is silent.

Without a sidecar, everything is meaningless

A Pod created before you attached istio-injection=enabled to the namespace has no sidecar. That Pod is Running and the service responds normally. But because traffic does not pass through the mesh, neither mTLS, nor authorization policy, nor routing is applied.

It is a state in which you have written all the security policies and nothing is being protected.

The way to check has changed

Istio now uses Kubernetes native sidecars. istio-proxy goes into spec.initContainers rather than spec.containers (with restartPolicy: Always).

So the habit of checking with kubectl get pod -o jsonpath='{.spec.containers[*].name}' now says there is no sidecar even when there is one. If you trust that answer, you judge that injection did not happen and end up fixing the wrong place.

The reason for the move was an ordering problem. In the old approach, the app started before the sidecar and sent requests before the network was ready, or the app finished but the sidecar remained so a Job never ended. A native sidecar starts before the app and ends after it.

Two kinds of "it doesn't work"

Symptom Cause Where to look
No response code (curl 000, exit 56) mTLS not satisfied — cut off at the TLS handshake Is the peer outside the mesh, and does it have a sidecar?
403 (RBAC: access denied) Authorization denial — connection and mTLS succeeded The AuthorizationPolicy's principals

A single response code splits the connection layer from the policy layer. Without this distinction, every time you look at a mesh failure you end up scanning from the beginning.

What a mesh actually does and its cost

Once a sidecar is attached, the path a single request travels gets longer. App → its own proxy → the peer's proxy → the peer app. This structure gets you mTLS, retries, circuit breakers, fine-grained routing, and metrics for every call without changing app code. In exchange there is a price to pay, and you must adopt it knowing that.

Latency is added. Because it passes through a proxy twice, milliseconds are added to each call. A single call is negligible, but if one screen makes twenty internal calls, it piles up accordingly.

It consumes resources. Because one proxy is attached per Pod, as much extra memory and CPU is needed as there are Pods. With hundreds of Pods, this total amounts to several nodes' worth.

One more point of failure is added. If a proxy cannot receive its configuration, traffic stops. Even if the control plane dies, it holds out for a while on the configuration it already received, but changes in the meantime are not reflected.

So set the criterion for judgment like this. If there are only a few services and the call relationships are simple, a mesh is an excessive choice. You can attach retries and circuit breakers with a library, and mTLS can be terminated at the gateway. A mesh earns its keep when there are more than a few dozen services, several languages make it impossible to unify a library, and each team's deployment cadence differs so that it is hard to enforce common policy in code.

When you adopt it too, do not turn everything on at once. Turn on injection starting with one namespace, and start mTLS at PERMISSIVE. This mode accepts both encrypted and plaintext connections, so communication is not cut even when the inside and outside of the mesh are mixed. After confirming in the metrics that plaintext connections have dropped to 0, raise it to STRICT. If you skip this order and apply STRICT from the start, all communication with Pods that have not yet been injected is cut, and it becomes an outage whose cause you cannot see. Rules for external traffic coming in through the gateway and for traffic inside the mesh easily drift apart, so when you write a policy it is better to make clear each time which of the two you mean.

What really matters in practice

After attaching the label, always recreate the Pods. istio-injection=enabled applies only to Pods that will be created from now on. If you leave Pods that are already running as they are, they stay outside the mesh forever, and since the service responds normally even in that state, nobody notices.

When checking for a sidecar, look at initContainers too. Since the move to native sidecars, the habit of looking only at spec.containers says something that exists does not. You must also look at kubectl get pod -o jsonpath='{.spec.initContainers[*].name}'.

Sort the layer first by response code. If there is no code at all (curl 000), it is the connection layer, and if it is 403, it is the policy layer. Without this one branch, every time you look at a mesh failure you end up scanning from the beginning.

In the next lab, you confirm these things yourself on a real Istio.