TT Lab
Get started
Learn Learning paths Courses

Istio Deep Dive — Why It Flows That Way

A Sidecar's Config Arrives in Two Parts

Continue in TT Lab

In one line

What the sidecar container starts is not Envoy but pilot-agent. pilot-agent gathers the container arguments, environment variables and mesh configuration and writes one bootstrap, and that one sheet holds only two things: who this proxy is (node) and where it gets its configuration (xds-grpc). Everything else is received from istiod.

Why this was needed

Envoy starts from one configuration file. But you cannot bake a different configuration file in advance for each of thousands of proxies in a mesh. Every Pod has a different IP and a different name, and every time a service is added, all the files would have to be rewritten. So Istio split the configuration in two: a small bootstrap decided when the Pod starts, and everything else received from istiod after it is up.

Splitting it this way creates a new kind of failure. If the bootstrap is wrong, the proxy does not start at all (CrashLoop), and if the bootstrap is right but it cannot reach istiod, the proxy is up but not ready (Ready 0/2). To tell the two symptoms apart, you need to know where and from what the bootstrap is made.

How it works

The istio-proxy in the injection output starts with proxy sidecar --domain $(POD_NAMESPACE).svc.cluster.local …. $(POD_NAMESPACE) is not shell syntax but Kubernetes syntax, which the kubelet replaces with the environment variable of the same name. The environment variables are of two kinds. There are values decided at injection time, such as CA_ADDR and TRUST_DOMAIN, and values like the Pod name, namespace and service account, which can be known only once the Pod is up, so they are received with fieldRef.

pilot-agent gathers these and makes the bootstrap.

Place in the bootstrap Ingredients
node.id Four fields joined with tildes: the role (sidecar), the Pod IP, <파드이름>.<네임스페이스> and <네임스페이스>.svc.cluster.local (the placeholders are the Pod name and the namespace)
node.cluster <워크로드>.<네임스페이스> (the placeholders are the workload and the namespace)
node.metadata The ISTIO_META_* variables with the prefix removed (CLUSTER_ID, MESH_ID, WORKLOAD_NAME …)
static_resources.clusters Just xds-grpc. The discoveryAddress of the mesh configuration (default istiod.istio-system.svc:15012)
dynamic_resources Receives CDS and LDS over ADS and waits indefinitely for the first response (initial_fetch_timeout: 0s)

istiod is both the configuration server and the certificate authority. So in a default installation, CA_ADDR and discoveryAddress point to the same 15012. Identity proof starts from two volumes. A projected service account token with its audience narrowed to istio-ca proves "I am this service account", and the root certificate in the istio-ca-root-cert ConfigMap confirms that the other side is the real istiod.

What it looks like in the field

Only the sidecar is in CrashLoopBackOff. You commonly run into this after touching a cluster name with a custom bootstrap (sidecar.istio.io/bootstrapOverride) or an EnvoyFilter. If the first line of the log is Unknown gRPC client cluster, the name ADS points to and the static cluster name have drifted apart. This kind is caught before deployment with --mode validate.

The Pod gets stuck at 1/2 Ready. The proxy is up but 15021 returns 503. If Envoy stays in PRE_INITIALIZING, it has not received its first configuration from istiod — a network policy blocked 15012, it could not get a certificate because the token's audience did not match, or istiod is down. The admin port is open even in this state, so first look at control_plane.connected_state.

Configuration gets mixed up in a multi-cluster setup. If CLUSTER_ID is the same in two clusters, istiod cannot tell the two proxies apart and the endpoints get mixed up. The label is the key that chooses the configuration.

Official documentation: pilot-agent · Debugging Envoy and Istiod · Envoy bootstrap

What you will do in the next lab

You pull the arguments, environment variables and volumes out of the injection output, and write a bootstrap of the same shape by hand and validate it. You see Envoy reject it before it even starts when one cluster name is wrong, and confirm through the admin port that in a Pod with no istiod, the proxy waits without becoming ready.