TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

Profiles and mesh settings: two decisions that are expensive to undo

Continue in TT Lab

In one line

Istio installation splits into two layers, "choosing a bundle called a profile" and "writing values into the mesh-wide configuration," and both can be read in full with istioctl manifest generate before they ever reach the cluster.

Why read the installation in advance

When people first put up a mesh, the most common thing they do is copy and paste the single installation command line from the documentation. Behind that one line are nearly 40 objects, and some of them hold cluster-wide permissions. More important is the defaults that command plants along with it. The mesh-wide configuration is a value that every proxy injected afterward reads, so if you change it later, every workload in the mesh receives the configuration again. The cost of reverting is much greater than the cost of putting it up.

The harshest example is outboundTrafficPolicy.mode. With the default ALLOW_ANY, the proxy lets through calls going out to unregistered outside addresses. If you change it to REGISTRY_ONLY, the proxy rejects traffic to unregistered destinations. If nobody gave notice on the day you turn it on, the calls that were going out to the payment processor, the internal SMTP, and the external log collector are all cut at once. Values like this must be "agreed on before installing," not "checked after installing." However, this setting alone does not create a security boundary that controls every external communication of a Pod. To block even the communication paths that the proxy does not capture, separate network controls are needed. The Istio security best practices also distinguish settings that find missing service registrations from mandatory external communication security policies.

Two knobs — profile and values

A profile is a bundle that chooses what to bring up. The four commonly used as of 1.24 split like this.

Profile Control plane Gateway Ambient components
minimal istiod None None
default istiod ingress None
demo istiod ingress + egress None
ambient istiod None ztunnel + istio-cni-node

minimal is one set of gateway (Deployment, Service, HPA, PDB, Role, RoleBinding, ServiceAccount) removed entirely, and demo is that plus one more set for egress. When you want to change just one value rather than a bundle, you use --set values.…. For example, --set values.gateways.istio-ingressgateway.autoscaleMin=3 leaves the profile as it is and changes only the minReplicas of the rendered HPA from 1 to 3. If you mix the two knobs, you get a state of "I thought I changed the profile, but only one value changed."

The mesh-wide configuration is the third place. You put it in with --set meshConfig.…, and in the rendered output it does not go in as a separate CRD but as a YAML chunk in the mesh key of the ConfigMap istio in the istio-system namespace.

istioctl manifest generate --set profile=minimal \
  --set meshConfig.outboundTrafficPolicy.mode=REGISTRY_ONLY \
  --set meshConfig.trustDomain=lab.internal

Here trustDomain is the value that decides the front part of a workload identity (SPIFFE ID), so if you change it later, every authorization policy written against the already-issued identities goes out of step. What you write under defaultConfig are the default behaviors of each individual proxy — things like whether to hold the app container until the proxy is ready, and how long to wait on shutdown.

If you get a value wrong, you are caught before even reaching the cluster. If you give a value that is not in the enumeration, the render ends with unknown value ... for enum istio.mesh.v1alpha1.MeshConfig.OutboundTrafficPolicy.Mode, and if you give a profile name that does not exist, it says the profile file could not be found. One more thing — in 1.24 there is no flag called --profile at all. --set profile= is right, and if you get it wrong it ends with unknown flag.

What it looks like in the field

The first is the report "the installation is the same but it behaves differently on each cluster." The cause is usually that people installed with different --set options and that command was recorded nowhere. The habit of keeping the rendered output in a repository and comparing it prevents this. A manifest can be made without a cluster, so you can put installation changes into code review.

The second is a mesh-wide setting whose scope was not narrowed. If one team asks for more proxy threads and you raise the mesh-wide defaultConfig.concurrency, the memory usage of thousands of Pods goes up together. What you use in such a case is the ProxyConfig CRD — it overrides the global default at the namespace or workload selector level. Value checking is done by the CRD schema, not istioctl, so a value such as a negative number is rejected by the API server.

Limits of this lab environment

In the lab Pod, neither istiod nor a gateway comes up. So you cannot directly see "I changed it to REGISTRY_ONLY and the calls were cut," and you cannot confirm the scene of the proxy actually reading this setting. Instead, the very document that decides the installation can be made completely offline, and since a real API server is attached, for settings that go in as a CRD like ProxyConfig you can actually confirm even the schema enforcement.

What you will do in the next lab

You render the four profiles in turn and compare in files what grows and shrinks, put in four mesh-wide settings, and find them in the rendered ConfigMap istio. You get rejected three times on purpose with a wrong value, a nonexistent profile, and a nonexistent flag, and at the end you write a script that builds a table of profiles and components and re-renders it yourself to verify.