A Zero-Trust Mesh — Identity, Encryption and Authorization
In one line
Mesh security has three layers. Who are you (the SPIFFE identity) → is encryption enforced (PeerAuthentication) → and so what can you do (AuthorizationPolicy). If you skip the order, an incident is certain.
Why this was needed
The inside of a Kubernetes cluster is often treated as a "trusted network," but in reality it is frequently a flat network where, if even one Pod is compromised, it can connect in plaintext to every service in the cluster. To block this in the application, you would have to put TLS certificate issuance, renewal, and verification logic and permission-check logic into every service. If there are four languages, there are four implementations, and the four behave subtly differently.
And there is a more fundamental problem. Network policy alone cannot express "who." IP-based control collapses when a Pod restarts, and it cannot stop impersonation by another Pod on the same node either. A mesh moves the basis of identity from the IP to the Kubernetes service account, and engraves that identity into a certificate and verifies it cryptographically.
How it works
Identity. The format follows the SPIFFE standard.
spiffe://cluster.local/ns/mesh-lab/sa/payments
------------- -------- --------
트러스트 도메인 네임스페이스 서비스 어카운트
istiod doubles as the CA. When a Pod comes up, istio-agent creates a key pair inside the Pod and sends only a CSR to istiod. The private key never leaves the Pod. istiod verifies the service account token, checks that the CSR's SPIFFE ID matches that token's namespace/SA, and then signs it. The certificate lifetime is 24 hours by default, and it is renewed automatically at about half its lifetime. The renewed certificate is delivered to Envoy through SDS, so neither a Pod restart nor connection draining is needed.
Enforcing encryption. PeerAuthentication decides "whether to require mTLS on incoming connections."
| Mode | Behavior | When to use |
|---|---|---|
| PERMISSIVE | Accepts both mTLS and plaintext (default) | The migration period |
| STRICT | Accepts only mTLS | The target state |
| DISABLE | Turns mTLS off | Exceptions, such as behind equipment that terminates TLS externally |
The scope has three layers and the narrower one wins — workload (selector specified) > namespace (that namespace without a selector) > mesh-wide (the name default in the root namespace istio-system). If you need an exception, open just one port with portLevelMtls.
You must not confuse the direction here. PeerAuthentication is a setting on the receiving side (server, inbound), and trafficPolicy.tls of a DestinationRule is a setting on the sending side (client, outbound). If the server is STRICT but the client-side DestinationRule is DISABLE, the connection fails entirely (the UF flag). istioctl analyze catches this conflict. The declaration on the client side that it will use a certificate issued by the mesh is ISTIO_MUTUAL.
End-user authentication. RequestAuthentication verifies a JWT's signature, issuer, and lifetime. There is a very common misunderstanding here. This resource alone blocks nothing. The rule is "if a token is present, it must be valid," and a request without a token just passes. To make a token mandatory, you must require requestPrincipals with an AuthorizationPolicy.
Authorization. The evaluation order is fixed for each request.
1. CUSTOM → 외부 인가기(OPA 등)가 거부하면 즉시 거부
2. DENY → 하나라도 매칭되면 즉시 거부
3. ALLOW → 대상 워크로드에 ALLOW 정책이 하나도 없으면 허용(기본 개방)
ALLOW 정책이 하나라도 있으면 매칭돼야 허용(기본 거부로 전환)
The last line is the key. The moment a single ALLOW policy exists, that workload goes into allowlist mode. The spec: {} idiom takes advantage of this property — an ALLOW policy exists but has no matching rule at all, so nobody gets through. This one line is the starting point of zero trust.
The value you write in principals is a SPIFFE ID, but remember that in configuration you write it without the spiffe:// prefix as cluster.local/ns/mesh-lab/sa/frontend. And a principal comes from the mTLS certificate, so a request that arrives in plaintext while in PERMISSIVE has an empty principal and never matches this condition. That is why moving to STRICT is a precondition for authorization design.
What it looks like in the field
First, STRICT that skipped the order. If you turn on mesh-wide STRICT without cleaning up the plaintext sources, crons without sidecars, legacy VMs, and custom probes are all cut off at once. The standard is observe (keep PERMISSIVE and check the plaintext ratio) → remove the plaintext sources → STRICT starting from non-critical namespaces → mesh-wide STRICT. The completion criterion is also set as a number — for example, "0 plaintext requests for 7 days."
Second, 403 in authorization. Three causes you often meet are (1) with PERMISSIVE the principal is empty and the match fails, (2) a typo in the namespace or service account, and (3) forgetting that the moment you add an ALLOW policy it switches to default deny, and not leaving the existing paths open.
Third, look at it in shadow before turning it on. The AUDIT action and the dry-run annotation do not block traffic and only record "what would have been denied if it were on." You must go through this stage before putting default deny into production.
Fourth, a dedicated service account per workload. If several workloads share the default SA, the SPIFFE ID becomes the same and you cannot split authorization. Identity design is authorization design.
What you will do in the next lab
Two labs follow. In the first lab, you start with PERMISSIVE and lay down a mesh-wide STRICT, create workload- and port-level exceptions, set ISTIO_MUTUAL on the client side, check what static analysis says when STRICT and DISABLE conflict, and leave a staged migration plan as a document. In the second lab, you start from an empty-rule full deny and narrow permissions by identity, method, path, namespace, and condition, layer on a DENY safety net and AUDIT, and then build a per-call allow/deny matrix.
In this environment no real mTLS handshake happens. Instead you judge the scope, priority, and conflicts of policies by manifests and static analysis. Most incidents in the field also come not from handshake failures but from setting the policy scope wrong.