TT Lab
Get started
Learn Learning paths Courses

System Integration (EAI)

Retries Are Not Free

Continue in TT Lab

Summary

When retries overlap across several layers they multiply and knock down the recovering counterpart again, so retry in only one layer and mix jitter into exponential backoff.

Why retries make outages bigger

The other system wobbles for a moment. We add a retry on our side. It seems like a good design. But consider this situation.

게이트웨이:   재시도 4회
  → 서비스A:  재시도 4회
    → 서비스B: 재시도 4회

The user pressed a button once, but 4 × 4 × 4 = 64 requests go to the final system. The counterpart that was recovering is knocked down again by this bombardment. And this wave repeats at every retry interval.

This is where the principle that you retry in only one layer comes from. If you retry in several layers, it amplifies exponentially.

Exponential backoff and jitter

Fixed-interval retries are bad in two ways. They hit again too soon, and they all hit together.

1차 실패 → 1초 대기 → 2차 → 2초 → 3차 → 4초 → 4차 → 8초 ... (상한 30초)

Exponential backoff solves the first problem. But the second remains.

If 10,000 clients follow the same rule, they retry at exactly the same moment. The recovering server is knocked down again at that moment. And that wave repeats every 2, 4, 8 seconds.

This is the thundering herd. The solution is jitter — mix randomness into the computed wait time.

sleep = random(계산값 × 0.5, 계산값)      # full jitter 계열

Backoff without jitter is only half done. And even without jitter, it works well normally, so it is discovered for the first time when a major outage occurs.

Errors that must not be retried

Without this distinction, retrying does harm.

Error type Retry? Reason
Connection refused, timeout Yes Likely to be temporary
HTTP 500, 502, 503, 504 Yes Temporary server-side error
HTTP 400 (format error) No Sent 100 times, fails 100 times
HTTP 401/403 (authentication/authorization) No Will not work until the credentials change
HTTP 404 No The target does not exist
HTTP 409 (duplicate) No It may mean it was already processed
HTTP 429 (limit exceeded) Conditional Respect Retry-After

Ignoring a 429 and retrying immediately is especially bad. The other side said "please come slowly," and you go faster. Many APIs respond to this with a block.

DLQ — you must know how to give up

A message that has exceeded the maximum attempts is sent to a DLQ (Dead Letter Queue). Infinite retry blocks the queue, and a blocked queue is a total stop.

When putting it in the DLQ, you must not put in only the original. Because a person has to look at it later.

{
  "msg_id": "M-20260819-000123",
  "original": { ...원본 전문... },
  "reason": "upstream 503 after 5 attempts",
  "attempts": 5,
  "first_failed_at": "2026-08-19T02:11:03+09:00",
  "last_failed_at": "2026-08-19T02:12:47+09:00",
  "trace_id": "a1b2c3d4"
}

And the DLQ is a monitoring target. If the count is not 0, someone has to look. There are really many systems where a DLQ is created and nobody looks at it. That is the same as silently throwing messages away.

Reprocessing design

How do you put back what has piled up in the DLQ?

  1. Classify — divide by reason. Is it a system error or a data error?
  2. Judge — fix it and put it back in, or ask the source to resend?
  3. Safety check — is reprocessing idempotent? If not, it becomes duplicate processing
  4. Execute — start with a small amount. If you put everything in at once, it gets blocked again for the same reason
  5. Record — when, who, and how many were reprocessed

Item 3 is the core. Reprocessing an interface that is not idempotent can produce a result "worse than not doing it." Duplicate deposits, duplicate purchase orders.

Idempotency — how to build it

Being idempotent means processing the same request many times causes the side effect only once. It does not mean "the response is always the same."

In the end there is one method. Attach a unique key to each request, and the receiver remembers the keys it has already processed.

CREATE TABLE inbox_log (
  msg_id     TEXT PRIMARY KEY,      -- ★ 멱등키. 제약이 마지막 방어선
  biz_key    TEXT NOT NULL,
  status     TEXT NOT NULL,
  response   TEXT,                  -- 최초 응답을 그대로 보관
  created_at TEXT NOT NULL
);

What to use as the idempotency key is the core of the design.

And always block it with a DB constraint (PRIMARY KEY / UNIQUE). An application-side if 조회 then 없으면 insert (check first, insert if absent) gets breached when two arrive at the same time. Because there is a gap between the check and the insert. A constraint removes that gap. Even if the application has a bug, the constraint cannot be breached.

Store and return

One step further, return the original response as is to a duplicate request.

1. msg_id 로 조회
2. 있으면 → 저장된 response 를 그대로 반환 (처리 안 함)
3. 없으면 → 처리하고, 결과를 response 에 저장

With this, resending becomes completely safe from the sender's point of view. When you could not receive a response because of a timeout, you can just send again. This is what the Idempotency-Key header convention does, and it is why payment APIs use this approach.

Retention period of the idempotency history

The inbox_log grows without limit. You must clean it up. But the moment you clean up, duplicate defense for that window disappears.

So set the retention period longer than "the maximum period in which a resend can realistically arrive." If the other side's reprocessing policy is "resend up to 7 days back," 30 days of retention is safe. And the cleanup batch deletes only by period condition. A method like "the oldest N first" deletes recent entries when inflow suddenly grows.

Four paths by which duplicates arise

Finally, let us sort out where duplicates come from in practice.

  1. Retry — resending after a timeout (when only the response was lost)
  2. The queue's at-least-once — the broker redelivers because it did not receive the ack
  3. An operator's manual resend — after an outage, "just send it again for now"
  4. A sender-side batch rerun — rerunning a failed batch from the beginning

Number 4 is the biggest. If a batch processed half and died and you rerun the whole thing, half is duplicated. So batches too must record a restart point or rely on receiver-side idempotency. Usually the latter is realistic.

What it looks like in the field

The form you see most often is the situation where nobody designed retries but retries are layered three deep. The gateway default, the HTTP client library default, and the code we wrote. Each is a reasonable 3–4 attempts, but multiplied it is dozens. So during outage response a conversation like "we sent only once, but the other side's log shows 40 entries" takes place.

The next most common thing is retrying errors that must not be retried. A format error or an authentication failure gives the same result even if you send it a hundred times, and if the retry logic does not distinguish status codes, those pointless retries fill the other side's logs and block our queue.

And at the end you must know how to give up. If you do not leave the reason, the number of attempts, and the first failure time together when moving to the DLQ, a person who opens that file weeks later cannot judge whether it is OK to put this back in.