When the Root Changes, Who Stops Trusting Whom
In one line
By default istiod signs workload certificates with a root it made itself. A plug-in CA is a way to bring the mesh's identity under the organization's PKI by plugging into istiod, through the cacerts Secret, an intermediate CA signed by the organization's root. The moment you change this on a mesh that is already running, workloads that trust only the old root and workloads that received new certificates cannot trust each other.
Why this was needed
A self-signed root is convenient to start with, but it has three problems. The root key is inside the cluster (istio-ca-secret), so if the cluster is breached, the root is breached too. The root differs per cluster, so workloads of two clusters cannot trust each other. And it is outside the organization's auditing and revocation system.
The structure the Istio docs recommend is this — keep the root CA on a highly secure offline machine, issue an intermediate CA per cluster from that root, and give it to each istiod. istiod signs workload certificates with that intermediate CA and distributes the root the administrator gave as the root of trust to the workloads. Even if one intermediate CA leaks, the root is intact, and clusters under the same root trust each other's workloads.
How it works
The four files of cacerts
| Key | Contents |
|---|---|
ca-cert.pem |
The intermediate CA certificate istiod will use |
ca-key.pem |
Its private key |
root-cert.pem |
The root to distribute to workloads |
cert-chain.pem |
The chain from the intermediate CA to the root |
istiod looks for istio-system/cacerts when it starts. If it is there, it uses it (the log notes that it read the root from cacerts), and if not, it uses the root it made itself. A running istiod does not notice a Secret created later, so you have to restart it.
The workload certificate is received by the sidecar. The sidecar's agent makes a key, asks istiod to sign it, receives a short-lived certificate (24 hours by default), puts it into Envoy through SDS, and swaps it on its own before it expires. The identity is carried not in the subject but in the SPIFFE URI of the SAN (spiffe://cluster.local/ns/bank/sa/client). An authorization policy's principals is exactly this value, so even if you change the CA, the policy stays the same as long as the trust domain is the same.
The moment you change the root. The new istiod changes istio-ca-root-cert in each namespace to the new root, and newly started Pods trust the new root and receive certificates signed by the new intermediate CA. But the sidecars that are already running hold only the old root. mTLS has both sides verify each other's chain, so a connection cannot be established between a web that received a new certificate and a client that trusts only the old root (503). This state continues until everything is restarted or the certificates are rotated.
The plug-in CA procedure in the Istio 1.31 docs covers the case of plugging in cacerts before installing. In this lab you see for yourself that changing it on a mesh that is already running produces the break above. The one conclusion from that observation is this — since the reason it breaks is 'they do not know each other's root', to change it in production you have to separate making every workload trust the new root from starting to sign with the new CA and set the order (a period of trusting both roots), and remove the old root only after everyone has received a new certificate.
What it looks like in the field
"I put in cacerts but the certificate issuer is unchanged." You did not restart istiod, or the Pod is still using the old certificate. Look at the istiod log and the sidecar's proxy-config secret separately.
"After changing the CA, only some service pairs are 503." These are pairs where only one side was restarted. If you look side by side at the two sidecars' ROOTCA and workload certificate issuer, it shows right away.
"I bound two clusters together but mTLS does not work between them." The two istiods are using different roots. You have to plug in, under the same root, an intermediate CA for each.
Official docs: Plug in CA Certificates · Security — PKI · Identity and certificate management
What you will do in the next lab
You first check the root istiod made itself, make a root and a cluster intermediate CA with the distribution's Makefile, and plug them in as cacerts. You restart only web to experience a 503 in the state where only one side received a new certificate, restart the client as well to revive it, and then read directly from the certificate whether the new chain leads up to our root, and read the identity and lifetimes.