Implementing Exponential Backoff, Jitter and a DLQ
Goal
You implement retries with exponential backoff and jitter, distinguish errors that must not be retried, and build DLQ loading, a reprocessing script, and a retry policy table.
Why it matters
Retries are not free. With gateway 4 attempts × service A 4 attempts × service B 4 attempts, one click by a user becomes 64 requests to the final system. The counterpart that was recovering is knocked down again by that bombardment. And exponential backoff without jitter makes 10,000 clients retry at exactly the same moment (the thundering herd). Since it is no problem normally, it is discovered for the first time when a really big outage happens. This lab treats retries not as "added or not added" but as design.
Steps
- Start the unstable API.
python3 /opt/lab/fixtures/eai/rest/flaky_api.py 9300(in the background)/flaky?key=<키>: for the same key, 503 through the 3rd call, 200 from the 4th (the placeholder is the key)/bad: always 400/dead: always 503
- Save the result of calling
/flaky?key=t1once to/root/r/first.txt. It must contain a linehttp_code=503. - Create
/root/r/retry.sh. It takes two arguments (URL 최대시도수, that is, the URL and the maximum number of attempts), retries at a fixed interval, and printsresult=<ok|fail> attempts=<n>on the last line. It exits with code 0 on success. - Create
/root/r/backoff.sh. It takes two arguments (URL 최대시도수, that is, the URL and the maximum number of attempts) and retries with exponential backoff (1, 2, 4, 8 seconds, capped at 30 seconds).- Before each attempt, print one line
attempt=<n> sleep=<초>(the placeholder is the seconds). - On the last line, print
result=<ok|fail> attempts=<n>. - If the environment variable is
DRY=1, do not actually wait and only print the plan.
- Before each attempt, print one line
- Add jitter to
backoff.sh. Thesleepvalue must be a random value between 50% and 100% of the computed value, and running twice withDRY=1must give different values. - Make
backoff.shdistinguish errors that must not be retried. If it receives HTTP 400/401/403/404/409, it must stop immediately and end withattempts=1. (Check with/bad) - Create
/root/r/senddlq.sh. It takes two arguments (URL 메시지ID, that is, the URL and the message ID), retries, and if it exceeds the maximum attempts it creates/root/r/dlq/<메시지ID>.json(the placeholder is the message ID). The JSON must have five keys:msg_id,url,reason,attempts, andfirst_failed_at. (urlis the target to call again in the step 7 reprocessing.) Run it with/deadand the message IDM-001to create the file. (If you run it again with the same ID, it overwrites the same file.) - Create
/root/r/dlq-replay.sh. It takes one argument (the DLQ directory), tries to reprocess each.jsonin it, and on failure increments that file'sattemptsby 1 and saves it again. On the last line, printreplayed=<시도수> succeeded=<성공수>(the number of attempts and the number of successes). If the directory does not exist, end with a non-zero exit code. - Create
/root/r/policy.csv. The first line iserror,retry,max_attempts,backoff,final. All six types below must be present.connection-refused,timeout,http-500,http-400,http-401,http-429.retryis one ofY/N/조건부(conditional), andfinalis one ofDLQ/중단(stop)/통보(notify).
Notes
- Getting only the status code:
curl -s -o /dev/null -w '%{http_code}' <URL> - bash random numbers:
$RANDOM(0–32767). Be careful with range conversion. - Building JSON:
jq -n --arg a "$A" '{msg_id:$a}'or directly with printf - Common mistake 1: the retry loop keeps going even after success. Break immediately on success.
- Common mistake 2: making the jitter larger than the computed value. If it exceeds the cap, the meaning of backoff blurs.
- Common mistake 3: naming the DLQ file with a timestamp, so files increase with every rerun. It must be keyed by message ID so that reprocessing can be managed.
Reproduce the failure
Save the result of calling /flaky?key=t1 once to /root/r/first.txt.
It must contain a line http_code=503.
Before building retries, you must first see what the failure looks like. Check the HTTP status code and the response body together.
Fixed-interval retry
Create /root/r/retry.sh. It takes two arguments (URL 최대시도수, that is, the URL and the maximum number of attempts),
retries at a fixed interval, and prints result=<ok|fail> attempts=<n> on the last line.
It exits with code 0 on success.
Taking the retry count as an argument makes it reusable. Exit immediately on success, and print how many attempts it took to succeed.
Exponential backoff
Create /root/r/backoff.sh. It takes two arguments (URL 최대시도수, that is, the URL and the maximum number of attempts) and
retries with exponential backoff (1, 2, 4, 8 seconds, capped at 30 seconds).
- Before each attempt, print one line
attempt=<n> sleep=<초>(the placeholder is the seconds). - On the last line, print
result=<ok|fail> attempts=<n>. - If the environment variable is
DRY=1, do not actually wait and only print the plan.
If it actually sleeps, grading becomes slow. Print the planned wait time first, and make it not sleep in dry-run mode.
Add jitter
Add jitter to backoff.sh. The sleep value must be
a random value between 50% and 100% of the computed value,
and running twice with DRY=1 must give different values.
If many clients follow the same rule, they retry at the same time. Mix randomness into the computed value, but it must not go out of range.
Distinguish errors that must not be retried
Make backoff.sh distinguish errors that must not be retried.
If it receives HTTP 400/401/403/404/409, it must stop immediately and
end with attempts=1. (Check with /bad)
A format error fails 100 times even if sent 100 times. Put a rule in the script that judges whether to retry by the status code.
DLQ loading
Create /root/r/senddlq.sh. It takes two arguments (URL 메시지ID, that is, the URL and the message ID),
retries, and if it exceeds the maximum attempts it creates /root/r/dlq/<메시지ID>.json (the placeholder is the message ID).
The JSON must have five keys: msg_id, url, reason, attempts, and first_failed_at.
(url is the target to call again in the step 7 reprocessing.)
Run it with /dead and the message ID M-001 to create the file.
(If you run it again with the same ID, it overwrites the same file.)
If you put only the original in the DLQ, a person cannot judge later. The reason, number of attempts, and first failure time must be together.
DLQ reprocessing script
Create /root/r/dlq-replay.sh. It takes one argument (the DLQ directory),
tries to reprocess each .json in it,
and on failure increments that file's attempts by 1 and saves it again.
On the last line, print replayed=<시도수> succeeded=<성공수> (the number of attempts and the number of successes).
If the directory does not exist, end with a non-zero exit code.
Reprocessing is safe only if it takes the target directory as an argument. If you do not leave a reprocessing history, you will not know how many times it was tried.
Retry policy table
Create /root/r/policy.csv. The first line is
error,retry,max_attempts,backoff,final.
All six types below must be present.
connection-refused, timeout, http-500, http-400, http-401, http-429.
retry is one of Y/N/조건부 (conditional), and final is one of DLQ/중단 (stop)/통보 (notify).
The retry decision and maximum count differ by error type. Be especially careful with 429, because it is the other side saying "come slowly."