The Moment Retries Make the Outage Bigger
In one line
A retry hides a transient failure. But if you leave retries on while the backend is overloaded, they multiply the load two or three times and complete the outage. That is why a retry always needs a cap and a budget.
Why this was needed
One backend gets slow → timeout → retry → load doubles → it gets slower → retry → …
This feedback loop is called a retry storm. Because a mesh puts a sidecar on every service, if each tier in a three-tier call chain retries 3 times, the service at the very end takes 27 times the load. Retries multiply.
There are three mechanisms that prevent it, and each solves a different problem.
| Mechanism | What it watches | What it does |
|---|---|---|
| Retry budget | The ratio of retries to total requests | Stops retrying once retries exceed a set ratio |
| Circuit breaker | Concurrent connections and pending requests | Makes requests fail immediately once a limit is exceeded (fails fast) |
| Outlier detection | Consecutive failures per endpoint | Takes bad endpoints out for a while |
How it works
Unlike its name, the circuit breaker is less a "breaker" than a limit.
# DestinationRule
trafficPolicy:
connectionPool:
tcp: { maxConnections: 100 }
http:
http1MaxPendingRequests: 20 # 커넥션을 기다리는 요청 상한
maxRequestsPerConnection: 100
A request over the limit does not wait and gets a 503 immediately. It looks harsh, but it is the right thing — if you make it wait, the client's threads get tied up too and the outage spreads upward. Failing fast is better for the whole system.
Outlier detection is the mechanism that takes the bad ones out of the load-balancing pool.
outlierDetection:
consecutive5xxErrors: 5
interval: 10s
baseEjectionTime: 30s
maxEjectionPercent: 50 # 절반 넘게는 절대 빼지 않는다
maxEjectionPercent is important. Without it, when the whole backend goes bad for a moment, every endpoint is taken out and the service dies entirely. The detection mechanism causes the outage.
Common misconceptions
"Retries are good to leave on" — if you put retries on a non-idempotent request (a payment or an order sent by POST), duplicates appear. Envoy's default retryOn holds only safe conditions, but the moment you add 5xx, a request that the server failed while processing is sent again too.
Retries without a timeout. With 3 retries and a 30-second timeout each, in the worst case you wait 90 seconds. By then the user has already left and only the connections remain tied up. You must set an overall time cap (timeout) together with the retries.
The safety device called a retry budget
What you always set together when you turn on retries is a retry budget. If you decide "use at most this percent of all requests on retries", the retry traffic does not run wild even when the downstream collapses.
# Envoy — 클러스터 단위 재시도 예산
circuit_breakers:
thresholds:
- priority: DEFAULT
max_retries: 3 # 동시에 진행 중인 재시도 상한
retry_policy:
retry_on: 5xx,reset,connect-failure
num_retries: 2
per_try_timeout: 1s
retry_back_off:
base_interval: 0.025s
max_interval: 1s
max_retries is a cap on the number of concurrent retries, so it physically prevents a retry storm. If you set only num_retries: 3 without it, the moment the downstream slows, one request becomes four and the load rises fourfold. You are sending four times the load to a service that is already collapsing.
per_try_timeout matters too. Without it, the first attempt uses up all the time within the overall timeout and there is no room left to retry. Overall timeout ≥ per_try × (retries + 1) must hold.
What may be retried and what may not
| Condition | Retry | Reason |
|---|---|---|
| Connection failure, TCP reset | ✅ | The server did not even receive the request |
| 503, 504 | ✅ | Usually rejected before processing |
| 500 | ⚠️ | It may have failed in the middle of processing |
| Timeout | ⚠️ | The server may have finished processing |
| 4xx | ❌ | Sending it again gives the same result |
A retry on timeout is the most dangerous. If a payment request times out at 5 seconds but the server completed it at 6 seconds, the retry means paying twice. So for non-idempotent requests, either remove timeout from retry_on, or always use an idempotency key.
Envoy lets you set this differently per request with the x-envoy-retry-on header, so you can be aggressive with reads and conservative with writes.
Removing a sick instance with outlier detection
Retries alone cannot fix the case where one specific instance is sick. That is because a retry can go to the same sick place. Outlier detection takes it out.
outlier_detection:
consecutive_5xx: 5 # 연속 5번 5xx 면
base_ejection_time: 30s # 30초 뺀다
max_ejection_percent: 50 # 다만 절반 이상은 못 뺀다
interval: 10s
max_ejection_percent is the safety device. When everything is sick at the same time (a shared DB outage), if you take them all out the service dies completely. Keeping half is better than collapsing.
What really matters in practice
In an outage investigation, the signals that tell these three apart are different.
# 서킷 브레이커에 걸렸나 — 대기열 넘침
envoy_cluster_upstream_rq_pending_overflow
# 이상치 감지가 엔드포인트를 뺐나
envoy_cluster_outlier_detection_ejections_active
# 재시도가 얼마나 도나
envoy_cluster_upstream_rq_retry / envoy_cluster_upstream_rq_total
If the last ratio exceeds 5%, retries are already part of the load. Increasing retries at that point is pouring oil on the fire.