The snack machine died before ACK
Do Not Overload the Snack Warehouse with Retries
In one line
A retry is not a button that erases failure. You decide which errors to retry, until when, and how many more times, and you must preserve that budget even after a restart.
Why this was needed
The festival snack order server slowed down for a moment. If all ten dispatchers start sending again immediately, the server gets retries piled on top of the existing work. Because the responses got slower, they retry again, and an outage that would have recovered in a moment gets longer. Even if the previous module prevented duplicate effects for the same business ID, the CPU and connections that process requests are not free. Preventing business duplicates and suppressing load are different problems.
Conversely, if you discard every failure as permanent, you lose accepted orders because of a brief connection error. That is why we divide failures into transient failures, permanent business errors, and failures whose meaning is unknown. In the lab, Retryable is an error that the transport adapter classified as retryable, and Permanent is an error whose content a person has to fix. An unknown exception is not packaged as a success or a quarantine but is propagated to the caller. We do not build a classifier this time that interprets a single HTTP status the same way for every service.
How it works
For one piece of work you store attempts, max_attempts, available_at, and deadline. attempts is not the number of network successes but the number of times a claim succeeded and an attempt was started. Even if the process goes down before sending the request, the attempt already consumed is not given back. This is because if you treat a failure whose send time is unknown as a free attempt, a job that dies at that same spot every time could repeat forever. This policy is a conservative budget that assumes failure detection and operator redrive.
The lab's delay formula is as follows. attempt starts at 1, and jitter is a ratio from 0 to 1 that the caller chooses. It rounds down to integer milliseconds and applies the cap first.
window = min(cap_ms, base_ms * 2**(attempt - 1))
delay_ms = floor(window * jitter)
With base=100, cap=1000, and jitter=0.5, you wait 50ms after the first failure and 100ms after the second. With base=100 and cap=250, the third attempt does not use 400 as is but 125ms, half of 250. The test feeds in 0, 1, and a fractional ratio directly to check the boundaries. If a real caller always uses only 0.5, every worker waits in the same pattern, so you do not get the effect of random spreading. The injectable ratio is a device for making the test deterministic, and it does not verify the random number generator itself.
available_at is a reservation that says it can be taken again only when this time comes. You do not sleep inside the function while holding the queue's write lock. You save the failure result as pending and finish the call, and a later worker takes the reservation that has arrived. A call with nothing to process ends with None. Here we implement a single unit of execution rather than building an always-on polling daemon, so production needs a separate wake-up, shutdown, and observation policy.
deadline is not extended at each new retry. Even if now+delay equals the deadline, there is no budget left to start at that point, so it is quarantined with the deadline reason. If it reached the maximum attempts, it is exhausted. If you have only exponential backoff and no count or deadline, it is just a program that repeats forever, slowly. Conversely, even if you set the count small, if a single send function never returns, you cannot finish the whole call. A time limit on the transport adapter is needed separately.
What it looks like in the field
When an outage happens, look at whether retries are duplicated across several layers. If the browser calls three times, the API server three times, and the external SDK three times, a single original piece of work can generate far more requests. Do not declare the load ceiling of the whole system from one layer's max_attempts alone. This lab's maximum attempts is the budget of a single jobs row, and it does not implement a service-wide per-second request limit or a token bucket.
It is also important that the time is stored in the DB. Time elapsed since process start can have a different baseline after a restart, so you do not use it as is for a persistent reservation time. In the lab, you pass integer ms that every worker interprets by the same baseline as an argument. run_once calls clock twice, before and after the send, and reports an error if the time goes backward within a single call. Clock differences across real hosts, reboots, and clock adjustments are separate design matters. This check does not sync the clocks of workers around the world.
When observing, look separately not only at the retry count but also at waiting work, leased work, quarantined work, and the age of the longest-waiting work. A shrinking pending count does not mean everything succeeded. It may have moved to dead. If another worker takes over after a lease expiry, the external request can be duplicated, so the same-ID handling of the receiving inbox you learned earlier is still needed. Keeping only one of the budget and duplicate prevention is not enough.
What you will do in the next check
In the quiz that follows, you calculate the delay cap, the deadline boundary, and the meaning of the attempt count. In the comprehensive lab after it, you implement retry_delay and a persistent queue and reject bool, infinity, and out-of-range values. You check that re-accepting with the same ID does not reset the existing deadline and the attempts already consumed.
Reference: AWS on backoff and jitter. The quantity and attempt ranges and the reservation and quarantine contracts are the design of this snack delivery lab.