TT Lab
Get started
Learn Learning paths Courses

Istio Service Mesh

Canary — The Abort Condition Matters More Than the Weight

Continue in TT Lab

In one line

A canary is not "give the new version a little traffic" but "first decide which metrics doing what means stop and roll back, and then give a little."

Why this was needed

Even a rolling update replaces Pods a little at a time. But what a rolling update looks at is only the health of the Pods. As long as the process is alive and the readiness probe passes, it keeps rolling. If the new version returns 200 nicely but the response content is wrong, p99 latency has tripled, or it returns 500 only for certain requests, a rolling update notices nothing and goes all the way to the end.

A mesh's traffic splitting adds one lever to this. It separates the number of Pods from the traffic ratio. You can bring up all the v2 Pods in advance and still give only 0% of the traffic, and if a problem shows up, you can return only the weight to 0 without touching the Pods. That rolling back requires no deployment is the decisive difference from a rolling update.

How it works

Weighted splitting is translated into Envoy's WeightedCluster. It draws a random number per request and picks a cluster according to the weight interval. So it is not exactly 90:10 but approaches 90:10 probabilistically. With only 100 requests, 12 may go to v2, and because of this property, for a service with little traffic you must wait at each canary stage until enough samples accumulate for the judgment to mean anything.

The sum of weights must be 100. If the sum is different, it is rejected at the validation stage, and thanks to this rule the accident of "lowering v1 to 90 but forgetting to raise v2" is structurally prevented.

You must tell three splitting methods apart and use them accordingly.

Method Target When to use
Weight A random portion of users General gradual rollout
Header match Only the designated people Early access for internal testers, QA
Mirroring Nobody (only a copy) Load and regression verification with zero user impact

Mirroring (shadowing) is especially misunderstood. The response to the copied request is thrown away entirely. Even if the mirror target is dead, there is no effect on the user's response. On the other hand, mirrored requests also really cause side effects such as DB writes and external API calls. To mirror a write path, preparation at the application level is needed so that the new version works in shadow mode. The Host header of a mirrored request gets a -shadow suffix, so you can tell them apart in logs, and you must.

There is one more trap in weighted splitting. Because a random number is drawn per request, the same user goes back and forth between v1 and v2 each time they move through screens. If session state or the UI differs between versions, this looks like a bug outright. The fix is the DestinationRule's consistent hash load balancing. It hashes a particular header (for example x-user-id) and always sends the same key to the same endpoint. Because it is a hash ring structure, remapping is minimal even as endpoints are added or removed.

Finally, a canary plan must be an executable definition, not a document. Each stage needs (1) a weight, (2) a criterion for moving to the next stage, and (3) an observation time, and the whole needs a rollback method. Flagger and Argo Rollouts take this definition as a CRD and run it automatically. The real value of automation is not speed but removing the human bias of hesitating to roll back even when seeing bad metrics.

What it looks like in the field

First, a canary that cannot even start because of a label mismatch. If there is not a single Pod with the label the subset points to, all the traffic that went to that subset is 503 (UH). The moment you raise the weight to 10, 10% of requests die. This is why many teams put a step in the deployment pipeline that verifies the subset labels match the Pod labels.

Second, a canary with no gate. Raising 10 → 30 → 50 → 100 by a person watching with their eyes is fine as a starting point, but night deployments are hard and judgment wavers. If you pin the error rate and latency thresholds into the stage definitions, that judgment leaves the person.

Third, a canary with nowhere to go back to. If you have already deleted v1, then turning the weight back to 0 leaves nowhere to go. The previous version must stay alive until the canary finishes and reaches 100%, and the time to clean it up must also be in the plan.

What you will do in the next lab

You start at 100/0 and move to 90/10, and check how a manifest whose sum is not 100 is rejected. You then attach in turn a route that sends only internal testers to the new version by header, mirroring that throws away the response, and consistent hashing that pins the same user to one version. At the end, you create a canary plan file containing the stages, the criteria, and the rollback, and advance the real mesh to the 50/50 middle stage.

There is one thing to note. The earlier steps of this lab grade saved files, not the cluster. vs-baseline.yaml, vs-90-10.yaml, vs-header.yaml, and vs-mirror.yaml must be kept separately as a snapshot of each step. If you keep rewriting one file, the evidence for earlier steps disappears. In practice too, committing a separate manifest for each canary stage is the standard — because which stage changed what becomes the material for an incident investigation.