TT Lab
Get started
Learn Learning paths Courses

Istio Deep Dive — Why It Flows That Way

Why the Circuit Breaker You Set Seems to Do Nothing

Continue in TT Lab

In one line

The trafficPolicy of a DestinationRule is carried over into fields of the Envoy cluster. loadBalancer becomes lb_policy, connectionPool becomes circuit_breakers.thresholds, and outlierDetection becomes outlier_detection. The carried-over numbers work as they are — a request that exceeds the limit does not even go to the upstream and becomes a 503, and an endpoint that failed consecutively is taken out after it has gone through the failures.

Why this was needed

Questions like "I set a circuit breaker but it has no effect" and "I turned on outlier detection but errors keep happening" come up endlessly. Most of the time it is not that the configuration is wrong but that the expectation is wrong. A circuit breaker limit is counted not for the whole service but separately for each cluster in each calling sidecar. Outlier detection is not prevention but an after-the-fact measure, so users take as many hits as the set count first. And if you give a subset its own policy, the parent policy is overwritten as a whole block, not field by field, so a limit you did not write quietly disappears.

None of these can be seen however long you stare at the YAML. You see them only by looking at how it was translated into the Envoy cluster, and actually overflowing it and making it fail.

How it works

DestinationRule Envoy cluster
loadBalancer.simple lb_policy
connectionPool.tcp.maxConnections circuit_breakers.thresholds[0].max_connections
connectionPool.http.http1MaxPendingRequests …max_pending_requests
connectionPool.http.http2MaxRequests …max_requests
outlierDetection.consecutive5xxErrors outlier_detection.consecutive_5xx
outlierDetection.baseEjectionTime · maxEjectionPercent base_ejection_time · max_ejection_percent

A limit you did not write becomes not Envoy's default of 1024 but the 4294967295 Istio puts in (effectively unlimited). A subset has its own cluster, so its policy is applied separately too, and the rule for combining them is one line in the official documentation — the subset's policy overrides the corresponding setting — and that is all. The unit of overriding is a block such as connectionPool, loadBalancer, outlierDetection or tls.

The circuit breaker's arithmetic is simple. If max_connections connections are tied up, new requests wait in a queue, and if the queue exceeds max_pending_requests, Envoy immediately returns a 503 and x-envoy-overloaded: true. Outlier detection counts consecutive failures per endpoint, and when it reaches consecutive_5xx, it takes that endpoint out of load balancing for base_ejection_time. But max_ejection_percent limits the share that can be taken out at once — because if you take them all out, there is nowhere left to send to.

What it looks like in the field

The circuit breaker does not open in a load test. The limit is counted in each calling sidecar. If there are ten calling Pods, the concurrent connections the service actually receives are ten times the limit. Conversely, a batch job with only one calling Pod hits a small limit right away.

Only the canary subset occasionally gets a flood of 503s. Either you wrote just tcp.maxConnections on the subset and the parent's http1MaxPendingRequests disappeared, or conversely you wrote a small http limit on the subset and forgot it. Put the subset cluster's limits side by side with the parent's using proxy-config cluster and look.

I turned on outlier detection but the 5xx on the dashboard does not reach 0. Failures before being taken out, and failures after the ejection time ends and it comes back, keep showing. Outlier detection is not a mechanism that removes errors but one that keeps it from dragging on for a long time. You have to use it together with retries for the errors users see to decrease.

Official documentation: DestinationRule · Circuit breaking · Envoy cluster

What you will do in the next lab

You write a DestinationRule with a traffic policy, build a field correspondence table and set up the cluster exactly as it says, then read it back with config_dump. You mix in an endpoint that always fails and count how many hits outlier detection takes before it removes it, confirm with /clusters that a subset policy overwrites the parent as a whole block, and pile concurrent requests onto a slow upstream to count the circuit breaker's 503s.