Implementing Exponential Backoff and a DLQ
Goal
Implement all four elements of retrying — exponential backoff, jitter, a cap, and a DLQ after giving up — and see it through to reprocessing, identifying the cause from only the information stored in the DLQ.
Why it matters
Retry code looks the easiest but causes incidents most often. Infinite retries at a fixed interval periodically knock down a server that is trying to recover. With only exponential backoff, 10,000 clients still retry at the same moment and create a wave. Even with jitter, without a cap you wait 17 minutes on the 10th attempt. And without a give-up condition, one bad message blocks the whole queue forever. This is called a poison message. A DLQ is a mechanism that isolates that message so the rest can flow, and at the same time it is a list of things to investigate. So putting in only the message is useless — the cause, the attempt count, and the first-seen time must be there together for you to fix the cause and reprocess.
Steps
- With
/root/qr/worker.py, build a worker that consumesq:tasks. The/opt/app/taskproc.pymodule'sprocess(msg)throws an exception whenmsg["kind"]isbad. - Put the failed message back in the queue with its attempt count incremented by 1. In
/root/qr/requeue.out,attempts=2must be visible. - In
/root/qr/schedule.txt, write the exponential backoff values for base 1 second and attempts 0–5 as 6 lines in the formattempt=<n> wait=<초>(wait in seconds). They are 1, 2, 4, 8, 16, 32. - In
/root/qr/jitter.txt, write 20 full-jitter wait times for attempt number 3, one per line. All must be between 0 and 8 inclusive, and at least 15 must be distinct values. - Apply a cap of 30 seconds and write the values for attempts 0–8 to
/root/qr/capped.txt. The last three values must all be 30. - Send a message that failed 5 times to
q:tasks:dlq.q:tasksmust be empty and the DLQ length must be 1. - The DLQ message JSON must have all four keys
payload,error,attempts, andfirst_seen_at, andattemptsis 5. /root/qr/replay.pytakes the message out of the DLQ, resetsattemptsto 0, and puts it intoq:tasks. After running it, the DLQ must be empty andq:tasksmust have 1 item.
Notes
- Full jitter:
wait = random.uniform(0, min(cap, base * 2 ** attempt)) - Errors to retry: timeouts, 503, 429. Do not retry: schema violations, nonexistent references.
- Common mistake 1: creating a DLQ consumer that reprocesses automatically — if the cause remains, it becomes an infinite loop.
- Common mistake 2: not setting an alert on the DLQ depth — a DLQ is quiet and is discovered weeks later.
Build a queue-consuming worker
With /root/qr/worker.py, build a worker that consumes q:tasks. The /opt/app/taskproc.py module's process(msg) throws an exception when msg["kind"] is bad.
Take out messages one at a time and pass them to the processing function. Keep the processing function able to fail on purpose.
Requeue the failed message
Put the failed message back in the queue with its attempt count incremented by 1. In /root/qr/requeue.out, attempts=2 must be visible.
If you keep the attempt count inside the message, you can continue counting on the next consumption.
Calculate the exponential backoff schedule
In /root/qr/schedule.txt, write the exponential backoff values for base 1 second and attempts 0–5 as 6 lines in the form attempt=<n> wait=<초> (wait in seconds). They are 1, 2, 4, 8, 16, 32.
Raise the attempt number as an exponent. Do just the calculation first, save it to a file to check, and then put it in the code.
Apply full jitter
In /root/qr/jitter.txt, write 20 full-jitter wait times for attempt number 3, one per line. All must be between 0 and 8 inclusive, and at least 15 must be distinct values.
It is a random number between 0 and the computed value. The value must differ each time even for the same attempt number.
Put a cap on the wait time
Apply a cap of 30 seconds and write the values for attempts 0–8 to /root/qr/capped.txt. The last three values must all be 30.
The exponent grows quickly. Without a cap, you wait 17 minutes on the 10th attempt.
Send to the DLQ after 5 failures
Send a message that failed 5 times to q:tasks:dlq. q:tasks must be empty and the DLQ length must be 1.
It is a retry only if there is a give-up condition. It must disappear from the original queue and exist only in the DLQ.
Put investigation information in the DLQ message
The DLQ message JSON must have all four keys payload, error, attempts, and first_seen_at, and attempts is 5.
You need four things to be able to investigate later: the original message, the failure cause, the attempt count, and the first-received time.
Reprocess after removing the cause
/root/qr/replay.py takes the message out of the DLQ, resets attempts to 0, and puts it into q:tasks. After running it, the DLQ must be empty and q:tasks must have 1 item.
Take it out of the DLQ, return it to the original queue, and reset the attempt count. It must be an explicit run, not an automatic loop.