Readiness Probes and Outlier Detection Watch Different Things
In one line
Kubernetes pulls traffic away based on the state a Pod reports about itself (the readiness check). The mesh's outlier detection pulls it away based on the responses the sidecar actually received. A Pod that passes the readiness check but returns a 500 on every request looks fine to the former eye and broken only to the latter.
Why this was needed
A readiness check usually looks at a light path such as /healthz. But outages often happen outside that path — the database connection pool dried up, or a single Pod with one setting gone wrong returns a 500 for certain requests. The readiness check is still 200, so Kubernetes keeps that Pod in the endpoints. With three Pods, a third of the requests fail, and it stays that way until someone looks at the logs.
The client-side proxy receives the responses directly, so it knows this. If the same upstream fails repeatedly, you only need to pull that upstream out for a while — Envoy's outlier detection does that job, and Istio turns it on with the DestinationRule's outlierDetection.
How it works
Outlier detection — it looks at each upstream host separately.
| DestinationRule | Envoy cluster | Meaning |
|---|---|---|
consecutive5xxErrors: 2 |
consecutive5xx: 2 |
Eject when there are this many consecutive 5xx |
interval: 2s |
interval: "2s" |
The period at which it judges whether to eject |
baseEjectionTime: 60s |
baseEjectionTime: "60s" |
The base time for one ejection |
maxEjectionPercent: 50 |
maxEjectionPercent: 50 |
The maximum ratio of hosts that can be pulled at once |
Ejection happens only in that sidecar's load-balancing table. The Kubernetes endpoints stay as they are, and the sidecars of other clients judge separately from the responses they received. An ejected host returns after the base ejection time passes, and if it fails again, it is out for longer in proportion to the number of ejections. maxEjectionPercent is a safety device that prevents pulling every host when the outage has spread to all hosts, leaving nowhere to go.
Making it visible. Istio filters out the Envoy statistics the sidecar produces by default. To see the ejection count (outlier_detection.ejections_enforced_total), you have to revive those statistics with proxyStatsMatcher in the Pod annotation proxy.istio.io/config, and since the proxy reads it when it starts, you must recreate the Pod. An ejected host shows OUTLIER CHECK as FAILED in istioctl proxy-config endpoints.
Mirroring — the mirror of a VirtualService copies the request and sends it to another service and discards that response. The copy goes with the same x-request-id as the original request, so you can find the pair in the receiving side's log. You test the new version with production traffic while returning only the original version's response to the user.
Fault injection — the fault of a VirtualService has the client-side sidecar add a delay (delay) or an abort (abort) before sending the request. An aborted request does not reach the upstream, and the response flag in the access log is printed as FI. If you attach a header condition, you can break only the request you are testing.
What it looks like in the field
"One of three Pods kept returning 500 and no one knew." It is a failure that passes the readiness check. If you put outlier detection on, the errors users see stop after a few. In exchange, the failure itself gets hidden, so you have to put an alert on the ejection statistics.
"I turned on outlier detection and all the hosts were pulled." It happens if you set maxEjectionPercent to 100 when the whole upstream got slow. If you pull everyone, there is nowhere to go, and the outage actually gets bigger.
"I turned on mirroring and orders went into the new version's database twice." A mirror only discards the response; the request is really processed. A request with side effects (a write) has to be separately excluded from the mirror target.
Official docs: DestinationRule — OutlierDetection · Mirroring · Fault Injection · Envoy outlier detection · ProxyConfig proxyStatsMatcher
What you will do in the next lab
You put one Pod that always returns 500 behind the same service as two healthy Pods and first measure the failure rate. After reviving the statistics, you put on outlier detection and confirm that the errors stop and that Envoy has ejected that Pod, and then add mirroring and header-conditioned fault injection. Finally you pull out which slot of the Envoy configuration these settings became.