TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

The hard part is not attaching the sidecar, it is living with it

Continue in TT Lab

In one line

The resources, startup order, and shutdown behavior of an injected proxy are decided by a few Pod annotations, and what fields those annotations actually become can be checked before putting anything on the cluster with istioctl kube-inject.

Why what comes after attaching is harder

Sidecar injection is easy to learn. Put a label on a namespace and each Pod gets one more container. The problem is what comes next. The proxy is not free, it must start and die together with the app, and for some workloads it gets in the way.

Let us start with resources. The default proxy requests CPU 100m and memory 128Mi, and its limits are 2 CPU cores and 1Gi of memory. Per Pod it is small, but with 2,000 Pods the requests alone reserve 200 CPU cores. Conversely, if you set the requests too low, the proxy slows down first when traffic surges, and you get an outage that looks like application latency. This knob must differ per workload, so you adjust it with Pod annotations.

The second is startup order. Containers in principle start at the same time. If the app is fast and the proxy is slow, requests going out in the app's first few seconds are caught only by the iptables rules without a proxy and simply fail. The 5xx that appear only right after deployment look like this. If you turn on holdApplicationUntilProxyStarts, the injector puts the proxy at the very front of the container list and attaches a postStart hook, holding back the start of the next container until the proxy is ready.

The third is the oldest trap. If the proxy goes into a Job Pod as an ordinary container, then even when the main work finishes, the proxy stays alive. The Pod cannot move on to completed, and the batch pipeline stops at that spot. In the past, people worked around it by calling the proxy's shutdown endpoint at the end of the job script — with the defect that if the job fails, that line is not executed. Now Kubernetes supports init containers that have a restart policy, so you move the proxy there. It starts first, stays alive while the Pod lives, and is cleaned up together when the main container finishes.

What field an annotation becomes

Annotation Where it changes in the injection result
sidecar.istio.io/proxyCPU, proxyMemory The resources.requests of istio-proxy
sidecar.istio.io/proxyCPULimit, proxyMemoryLimit The resources.limits of istio-proxy
proxy.istio.io/config The environment variable PROXY_CONFIG of istio-proxy (converted to JSON)
holdApplicationUntilProxyStarts in proxy.istio.io/config Container order + a lifecycle.postStart hook
traffic.sidecar.istio.io/excludeOutboundPorts The command-line arguments of istio-init
sidecar.istio.io/inject: "false" The proxy and the init container are not included at all
sidecar.istio.io/nativeSidecar: "true" The proxy moves to initContainers and receives restartPolicy: Always

One important property here — proxy.istio.io/config has a string value but holds YAML inside it. That means the format a person writes and the format the proxy reads differ, so even if you make a typo, the manifest syntax still passes. It is why you need the habit of opening the resulting environment variable yourself.

The place you attach annotations is often wrong too. Injection targets the Pod, so the annotations must be attached to the Pod template. If you attach them to the Deployment's metadata, the syntax is right and nothing happens.

What it looks like in the field

The most common thing you see is the report "errors appear for just a few seconds right after deployment." Only connection refused remains in the logs and it cannot be reproduced. It disappears when you turn on the startup order guarantee. The reason this setting is not the default is that Pod startup gets slower by that much, so where to turn it on is decided by the nature of the service.

The second is "the nightly batch has not finished since yesterday." The Pod is Running and the application log printed a normal exit. Only the proxy is alive. If such Pods pile up, the batch queue gets blocked, and if nobody knows the cause, the practice of a person deleting Pods by hand becomes entrenched.

The third is database connections. If the proxy misidentifies the protocol (when the port name is missing or does not follow the rule), connections are cut or become strangely slow. The knob used temporarily then is excluding outbound ports — the root fix is to name the port to match the protocol, but during an outage this one annotation buys time.

Limits of this lab environment

In the lab Pod you cannot bring up a real proxy. So you cannot see the startup order guarantee actually saving the first request, a Job Pod failing to move on to completed, or how much memory the proxy uses. What this lab covers is what Pod spec a declaration becomes. Fortunately, the mistakes you can catch at this stage account for a large share of incidents in the field — attaching an annotation in the wrong place, getting the value format wrong, or something you thought you turned on not being on.

What you will do in the next lab

You first read the default injection result with no annotations, then add annotations one at a time and check which fields of the resulting manifest change. You look in turn at resources, proxy configuration, startup order guarantee, excluding outbound ports, and opting out of injection, and you inject a Job in two ways and compare which list the proxy goes to. At the end you build a script that re-injects eight manifests at once and hardens the results into a table.