TT Lab
Get started
Learn Learning paths Courses

Envoy Internals

The Moment Retries Make the Outage Bigger

Continue in TT Lab

In one line

A retry hides a transient failure. But if you leave retries on while the backend is overloaded, they multiply the load two or three times and complete the outage. That is why a retry always needs a cap and a budget.

Why this was needed

One backend gets slow → timeout → retry → load doubles → it gets slower → retry → …

This feedback loop is called a retry storm. Because a mesh puts a sidecar on every service, if each tier in a three-tier call chain retries 3 times, the service at the very end takes 27 times the load. Retries multiply.

There are three mechanisms that prevent it, and each solves a different problem.

Mechanism What it watches What it does
Retry budget The ratio of retries to total requests Stops retrying once retries exceed a set ratio
Circuit breaker Concurrent connections and pending requests Makes requests fail immediately once a limit is exceeded (fails fast)
Outlier detection Consecutive failures per endpoint Takes bad endpoints out for a while

How it works

Unlike its name, the circuit breaker is less a "breaker" than a limit.

# DestinationRule
trafficPolicy:
  connectionPool:
    tcp: { maxConnections: 100 }
    http:
      http1MaxPendingRequests: 20     # 커넥션을 기다리는 요청 상한
      maxRequestsPerConnection: 100

A request over the limit does not wait and gets a 503 immediately. It looks harsh, but it is the right thing — if you make it wait, the client's threads get tied up too and the outage spreads upward. Failing fast is better for the whole system.

Outlier detection is the mechanism that takes the bad ones out of the load-balancing pool.

outlierDetection:
  consecutive5xxErrors: 5
  interval: 10s
  baseEjectionTime: 30s
  maxEjectionPercent: 50      # 절반 넘게는 절대 빼지 않는다

maxEjectionPercent is important. Without it, when the whole backend goes bad for a moment, every endpoint is taken out and the service dies entirely. The detection mechanism causes the outage.

Common misconceptions

"Retries are good to leave on" — if you put retries on a non-idempotent request (a payment or an order sent by POST), duplicates appear. Envoy's default retryOn holds only safe conditions, but the moment you add 5xx, a request that the server failed while processing is sent again too.

Retries without a timeout. With 3 retries and a 30-second timeout each, in the worst case you wait 90 seconds. By then the user has already left and only the connections remain tied up. You must set an overall time cap (timeout) together with the retries.

The safety device called a retry budget

What you always set together when you turn on retries is a retry budget. If you decide "use at most this percent of all requests on retries", the retry traffic does not run wild even when the downstream collapses.

# Envoy — 클러스터 단위 재시도 예산
circuit_breakers:
  thresholds:
    - priority: DEFAULT
      max_retries: 3          # 동시에 진행 중인 재시도 상한
retry_policy:
  retry_on: 5xx,reset,connect-failure
  num_retries: 2
  per_try_timeout: 1s
  retry_back_off:
    base_interval: 0.025s
    max_interval: 1s

max_retries is a cap on the number of concurrent retries, so it physically prevents a retry storm. If you set only num_retries: 3 without it, the moment the downstream slows, one request becomes four and the load rises fourfold. You are sending four times the load to a service that is already collapsing.

per_try_timeout matters too. Without it, the first attempt uses up all the time within the overall timeout and there is no room left to retry. Overall timeout ≥ per_try × (retries + 1) must hold.

What may be retried and what may not

Condition Retry Reason
Connection failure, TCP reset ✅ The server did not even receive the request
503, 504 ✅ Usually rejected before processing
500 ⚠️ It may have failed in the middle of processing
Timeout ⚠️ The server may have finished processing
4xx ❌ Sending it again gives the same result

A retry on timeout is the most dangerous. If a payment request times out at 5 seconds but the server completed it at 6 seconds, the retry means paying twice. So for non-idempotent requests, either remove timeout from retry_on, or always use an idempotency key.

Envoy lets you set this differently per request with the x-envoy-retry-on header, so you can be aggressive with reads and conservative with writes.

Removing a sick instance with outlier detection

Retries alone cannot fix the case where one specific instance is sick. That is because a retry can go to the same sick place. Outlier detection takes it out.

outlier_detection:
  consecutive_5xx: 5              # 연속 5번 5xx 면
  base_ejection_time: 30s         # 30초 뺀다
  max_ejection_percent: 50        # 다만 절반 이상은 못 뺀다
  interval: 10s

max_ejection_percent is the safety device. When everything is sick at the same time (a shared DB outage), if you take them all out the service dies completely. Keeping half is better than collapsing.

What really matters in practice

In an outage investigation, the signals that tell these three apart are different.

# 서킷 브레이커에 걸렸나 — 대기열 넘침
envoy_cluster_upstream_rq_pending_overflow

# 이상치 감지가 엔드포인트를 뺐나
envoy_cluster_outlier_detection_ejections_active

# 재시도가 얼마나 도나
envoy_cluster_upstream_rq_retry / envoy_cluster_upstream_rq_total

If the last ratio exceeds 5%, retries are already part of the load. Increasing retries at that point is pouring oil on the fire.