TT Lab
Get started
Learn Learning paths Courses

Queues and Asynchronous APIs

Implementing Exponential Backoff and a DLQ

Continue in TT Lab

Goal

Implement all four elements of retrying — exponential backoff, jitter, a cap, and a DLQ after giving up — and see it through to reprocessing, identifying the cause from only the information stored in the DLQ.

Why it matters

Retry code looks the easiest but causes incidents most often. Infinite retries at a fixed interval periodically knock down a server that is trying to recover. With only exponential backoff, 10,000 clients still retry at the same moment and create a wave. Even with jitter, without a cap you wait 17 minutes on the 10th attempt. And without a give-up condition, one bad message blocks the whole queue forever. This is called a poison message. A DLQ is a mechanism that isolates that message so the rest can flow, and at the same time it is a list of things to investigate. So putting in only the message is useless — the cause, the attempt count, and the first-seen time must be there together for you to fix the cause and reprocess.

Steps

  1. With /root/qr/worker.py, build a worker that consumes q:tasks. The /opt/app/taskproc.py module's process(msg) throws an exception when msg["kind"] is bad.
  2. Put the failed message back in the queue with its attempt count incremented by 1. In /root/qr/requeue.out, attempts=2 must be visible.
  3. In /root/qr/schedule.txt, write the exponential backoff values for base 1 second and attempts 0–5 as 6 lines in the form attempt=<n> wait=<초> (wait in seconds). They are 1, 2, 4, 8, 16, 32.
  4. In /root/qr/jitter.txt, write 20 full-jitter wait times for attempt number 3, one per line. All must be between 0 and 8 inclusive, and at least 15 must be distinct values.
  5. Apply a cap of 30 seconds and write the values for attempts 0–8 to /root/qr/capped.txt. The last three values must all be 30.
  6. Send a message that failed 5 times to q:tasks:dlq. q:tasks must be empty and the DLQ length must be 1.
  7. The DLQ message JSON must have all four keys payload, error, attempts, and first_seen_at, and attempts is 5.
  8. /root/qr/replay.py takes the message out of the DLQ, resets attempts to 0, and puts it into q:tasks. After running it, the DLQ must be empty and q:tasks must have 1 item.

Notes

Build a queue-consuming worker

With /root/qr/worker.py, build a worker that consumes q:tasks. The /opt/app/taskproc.py module's process(msg) throws an exception when msg["kind"] is bad.

Take out messages one at a time and pass them to the processing function. Keep the processing function able to fail on purpose.

Requeue the failed message

Put the failed message back in the queue with its attempt count incremented by 1. In /root/qr/requeue.out, attempts=2 must be visible.

If you keep the attempt count inside the message, you can continue counting on the next consumption.

Calculate the exponential backoff schedule

In /root/qr/schedule.txt, write the exponential backoff values for base 1 second and attempts 0–5 as 6 lines in the form attempt=<n> wait=<초> (wait in seconds). They are 1, 2, 4, 8, 16, 32.

Raise the attempt number as an exponent. Do just the calculation first, save it to a file to check, and then put it in the code.

Apply full jitter

In /root/qr/jitter.txt, write 20 full-jitter wait times for attempt number 3, one per line. All must be between 0 and 8 inclusive, and at least 15 must be distinct values.

It is a random number between 0 and the computed value. The value must differ each time even for the same attempt number.

Put a cap on the wait time

Apply a cap of 30 seconds and write the values for attempts 0–8 to /root/qr/capped.txt. The last three values must all be 30.

The exponent grows quickly. Without a cap, you wait 17 minutes on the 10th attempt.

Send to the DLQ after 5 failures

Send a message that failed 5 times to q:tasks:dlq. q:tasks must be empty and the DLQ length must be 1.

It is a retry only if there is a give-up condition. It must disappear from the original queue and exist only in the DLQ.

Put investigation information in the DLQ message

The DLQ message JSON must have all four keys payload, error, attempts, and first_seen_at, and attempts is 5.

You need four things to be able to investigate later: the original message, the failure cause, the attempt count, and the first-received time.

Reprocess after removing the cause

/root/qr/replay.py takes the message out of the DLQ, resets attempts to 0, and puts it into q:tasks. After running it, the DLQ must be empty and q:tasks must have 1 item.

Take it out of the DLQ, return it to the original queue, and reset the attempt count. It must be an explicit run, not an automatic loop.