TT Lab
Get started
Learn Learning paths Courses

ICA — Istio Certified Associate

Where to Send It, and How to Treat It on Arrival

Continue in TT Lab

In one line

A VirtualService decides where to send, and a DestinationRule decides how to treat it once it arrives. Half of the incidents where canary deployments do not work arise from building only half of this division of labor.

Why this was needed

The starting point is the wish to separate deployment from release. If you can control separately putting a new version on the cluster (deployment) and sending user traffic to it (release), then when a problem occurs, instead of rolling back the image you only have to lower the weight to 0.

But these two are different in nature. A distribution such as "only 10% to v2" is a decision per request, while the definition that "v2 is the Pods with the label version=v2" is a fact per destination. Istio split this into two resources.

This produces an asymmetry that comes up on the exam as is. A subset exists only in a DestinationRule, and a weight exists only in a VirtualService. If you reference subset: v2 in a VirtualService without a DestinationRule, no such cluster is created in Envoy, so the request gets a 503. The response flag in the access log is recorded as UH (no healthy upstream).

How it works

Routing rules are evaluated from top to bottom, and the first match wins. So you put specific rules at the top and the catch-all rule at the very bottom. If there is no final rule without a match (catch-all), a request that matches no rule becomes a 404 or 503 with the NR (no route) flag.

The combination rule for match conditions is also often gotten wrong.

# AND: 경로가 /admin 이면서 동시에 헤더도 맞아야 매칭
- match:
    - uri:
        prefix: /admin
      headers:
        x-role:
          exact: admin

# OR: 경로가 /admin 이거나, 헤더가 맞으면 매칭
- match:
    - uri:
        prefix: /admin
    - headers:
        x-role:
          exact: admin

Conditions inside one match block are AND, and between items of the match array is OR. A one-space difference in indentation flips the meaning of the policy.

Resilience settings should be read by separating the time budget from the retry conditions rather than by memorizing fixed multipliers. attempts: 2 is not two attempts including the original request but at most two additional retries, that is, at most three upstream requests. perTryTimeout is the time cap for the original request and each retry, and timeout is the limit for the whole request including retries and waiting. It does not mean you always wait the full cap on every attempt.

Even when the overall limit and the per-try limit are both 1 second, fast failures are different. If you specify retryOn: "503" and the actual upstream returns a 503 immediately, retries can be executed within the overall 1 second. On the other hand, if the first attempt uses up the entire budget without a response, there is no time left for the next attempt. So from the fact alone that the overall limit is smaller than perTryTimeout × (attempts + 1), you must not conclude that retries are impossible. If the budget is meant to give every attempt time up to its cap, you must consider not only that product but also the backoff between attempts.

The actual count varies with the failure condition, the timing of the failure, the backoff, and the remaining budget. In a separate measurement on Istio 1.31.0, with the same 1 second / 1 second / 2 additional retries setting, a fast 503 reached the upstream three times, and a response taking 1.5 seconds reached it once. Do not guess the count from the final HTTP code; cross-check the request identifier against the upstream's receipt records. A case where the proxy turns a connection error into a 503 differs from an actual upstream 503, so retryOn: "503" alone does not guarantee the same retries.

Decide the per-try limit by looking at the normal latency distribution and the caller's remaining budget, and turn on retries only after confirming that the business operation is safe against duplicate processing. The name GET alone does not guarantee that an implementation has no side effects, and there is no rule that a single retry of a POST is safe. Even if you did not receive a response, the server may have already written. In the synthetic order measurement, three writes remained after a single success response, while in the control group that atomically deduplicated the same request key, only one remained. Do not interpret the single-process in-memory control group as a production exactly-once guarantee.

For the official definition, check HTTPRetry in VirtualService. The file-writing lab below practices the configuration structure, and the HTTP observations in this explanation were made on a separate VM with a real sidecar.

Retries must be viewed from the perspective of the whole chain. In A → B → C, if each hop has attempts 2, the requests C receives become 3 x 3 = 9 times in the worst case. With four hops, it is 27 times. Since retries amount to nailing the coffin shut on a dying service, put retries in only one layer of the chain (preferably the outermost).

The circuit breaker is gathered not in two resources but in one place, the DestinationRule. connectionPool decides "how much to let pile up," and outlierDetection decides "which instances to pull out." If you set maxEjectionPercent: 100 on a service with only three instances, all three are ejected at the same time and the circuit breaker itself becomes a total outage.

The same DestinationRule also has localityLbSetting. It is the setting that uses endpoints in the same zone first and, if that zone dies, moves on to the next zone. Traffic that crosses zones carries cost and latency, so at scale you inevitably turn it on.

There is a trap here that shows up both on the exam and in the field. Failover does not happen without outlierDetection. The condition for moving on to the next priority is "the endpoints of the current priority are unhealthy," and the mechanism that produces that judgment is outlierDetection. Without it, Envoy has no way of knowing that the endpoints in the same zone are dying, and it keeps sending there.

The reason the symptom is confusing is that it works fine in normal times. The zone-preferred routing itself works well, latency drops, and cost goes down. Then it fails to move over only on the day one zone dies. If the label (topology.kubernetes.io/zone) were wrong, the zone-preferred routing would not work in the first place, so the symptom differs — if it works well normally but fails only during an outage, it is almost always that outlierDetection is missing.

What it looks like in the field

I am carrying over as is one lesson from the author's homelab. On that Cilium-based cluster, when a rules.http: [{method: GET}] policy was applied to an nginx with the same IP and the same port, GET returned 200 and POST returned 403. Since the method is not visible at the L3/L4 layer, this distinction means someone parsed HTTP. In Istio that someone is the sidecar Envoy, and the reason header-matching routing and method matching are possible is exactly the same.

What was interesting was the shape of the Hubble log at that time. The request was recorded as DROPPED and the response as FORWARDED. This is because the 403 created by the proxy flowed out as a normal response. In Istio too, the situation where the connection is established but a 403 comes back and the situation where the connection itself fails have different log shapes. The former is an L7 judgment and the latter is an L4 problem. If you cannot make this distinction, you end up staring at the wrong resources for hours.

What you will do in the next lab

Two labs follow. In the first, after actually deploying the namespace and the v1/v2 workloads, you write the DestinationRule and VirtualService manifests as files and get hands-on with weights, header matching, URI rewriting, and rule order. In the second, you cover the relationship between retries and timeouts, fault injection gated by a header, circuit breakers, and mirroring.