What Infinite Retries Create
Summary
Retries need a cap, exponential backoff, jitter, and a place to go after giving up. If even one of the four is missing, retrying is an outage amplifier.
Why this was needed
The worst retry is one done immediately, at a fixed interval, and without limit. When a server briefly slows down from overload, if every client immediately hits it again, the server that was barely holding on collapses completely. Retries make the outage worse instead of curing it.
The first improvement is exponential backoff. You lengthen the interval to 1 second, 2 seconds, 4 seconds, 8 seconds. As failures continue, retry pressure decreases exponentially, giving the server room to breathe.
The second is jitter. Exponential backoff alone is not enough. If the server dies briefly, 10,000 clients detect the failure at the same time, and all follow the same rule and retry at exactly the same moment. Just as the server is trying to recover, 10,000 requests pour in and it collapses again. And this wave repeats after 2 seconds, 4 seconds, and 8 seconds. Full jitter breaks this synchronization by waiting a random time between 0 and the computed backoff value.
How it works
You also need a cap and giving up. You set a maximum wait time (for example, 30 seconds) so the interval does not grow indefinitely, and give up after a certain number of attempts. The place a message that gave up goes is the DLQ (dead-letter queue).
A DLQ is not a trash can but a list of things to investigate. So putting in only the message is useless. There are four things to put in together — the original message, the failure cause (an exception message or status code), the number of attempts, and the time it was first received. With these, you can answer "why is this message here?", and you can reprocess it after fixing the cause.
You should also prepare a reprocessing script in advance. If the DLQ piles up at 3 a.m. and you build the reprocessing tool on the spot, mistakes happen.
And you must not retry just any failure. Retrying is meaningful only for transient errors. When the message itself is wrong (a schema violation, a nonexistent reference), it is the same after 100 attempts. It is better to send these straight to the DLQ on the first failure. If you do not make this distinction, one bad message blocks the whole queue — this is called a poison message.
What you meet in the field
Many teams do not set an alert on the DLQ. Because a DLQ is quiet, it is discovered weeks later. "DLQ depth > 0" is a metric that, even if it is not a page, someone should at least look at once a day.
One more thing. A design that creates a consumer of the DLQ itself and lets it reprocess automatically is dangerous. If the cause remains, it becomes an infinite loop. It is safer for reprocessing to be triggered explicitly after a person has checked the cause.
When backoff has no jitter
Failed jobs retry clustered at the same moment. The moment a downstream service dies briefly and comes back, that crowd arrives all at once and kills it again (thundering herd).
# ❌ 모두가 정확히 1s, 2s, 4s 뒤에 재시도한다
delay = base * (2 ** attempt)
# ✅ 흩어진다 — full jitter
delay = random.uniform(0, base * (2 ** attempt))
# ✅ 최소 대기는 보장하면서 흩기 — decorrelated jitter
delay = min(cap, random.uniform(base, prev * 3))
full jitter is the simplest and in measurements scatters well too. Put a cap (cap) so that the exponent does not grow without limit.
What decides the number of retries
"Three times" is merely a convention. It is better to calculate backward from the total wait time.
하류가 보통 30초 안에 복구된다면 → 총 대기가 60초쯤 되게 잡는다
base=1s, 지수 2, 상한 20s → 1, 2, 4, 8, 16, 20 … 6번이면 51초
And you give up immediately on failures where retrying is useless. 400 (bad request), 401 and 403 (permission), and 404 are the same no matter how many times you send them. If you do not distinguish these, a job that will fail forever blocks the queue.
class Permanent(Exception): pass # 재시도하지 않는다
if 400 <= status < 500 and status not in (408, 429):
raise Permanent(f"고칠 수 없는 실패: {status}")
How to operate a dead-letter queue
Jobs that have used up their retries are sent to the DLQ. But if you stop at putting them in the DLQ, nobody looks. You need three things together.
- Store the reason together — the last error, the number of attempts, the original message, the time of first failure.
- Alert on the count — if something that was usually 0 grows, that is itself a signal.
- A path to put them back — there must be a command that returns items from the DLQ to the original queue after the cause is fixed. If you leave people to copy by hand, nobody does it.
# 예: 원인을 고친 뒤 되돌리기
labhub-cli dlq replay --queue orders --since 2026-09-06T12:00 --limit 500
It is important not to put them back all at once. If you re-inject 5,000 items at once, that itself becomes a spike and collapses things again. Apply a rate limit and feed them in gradually.
What you will do in the next lab
You attach exponential backoff and full jitter to a queue consumer, set a cap, send to the DLQ after 5 failures, put the cause and the attempt count in the DLQ message, and finally remove the cause and recover with a reprocessing script.