The snack machine died before ACK
Do Not Erase Failed Snack Orders
In one line
Quarantine is not turning a failure into a success. You leave the evidence of the failure, stop the automatic repetition, and only a redrive that a person allowed starts with a new budget.
Why this was needed
One snack order keeps having the wrong content. If you repeat the same work immediately, it takes the slot for processing the normal orders behind it. But if you delete the row, you cannot explain where the order the customer placed went. This time we split work into done and dead and preserve the reason it stopped. Growth in dead is not a rise in the normal completion rate; it is an increase in work that needs investigation.
The outbox delivery loop in the earlier module stopped at the first failure. This queue deals with independent snack orders, so it can process other work whose reservation time has passed. This choice does not fit every event stream. You must not apply it the same way to work where changing the order changes the meaning, such as bank balance calculation or document editing. You must first decide how quarantining a failed item affects the ordering guarantee.
How it works
finish takes the transport adapter's result as ok, retry, or permanent. ok is done and permanent is dead. Even for retry, if it reached the maximum attempts it becomes dead with exhausted, and if the time it wants to schedule reaches the original deadline it becomes dead with the deadline reason. If budget remains, it stores pending and the next available_at. On no path does it delete the business ID or quantity; it releases only owner and lease_until.
run_once classifies it as ok only when the network adapter returned an exact True. A 1 or the string ok is not considered a success. You state a transient failure exception explicitly as Retryable and a permanent failure as Permanent. An unexpected exception or a BadAck is propagated to the caller and leaves the leased state that has not been finalized. The next claim checks the lease expiry and reclaims it, so instead of hiding the failure, the basis for recovery is kept.
To run a quarantined work item again, you explicitly call redrive. It is not allowed for pending, leased, or done. The note an operator leaves is an identifier pointing to a diagnostic ticket or a fix task. In the lab, it accepts only an ASCII identifier so that raw error messages or personal information do not slip into the audit record. A real service would also need who approved it and access permissions, but this local function does not implement that authentication scheme.
The redrive does adding an audit row and updating the work state in one transaction. In redrives it records id, the previous token, at_ms, and note. It resets jobs' attempts to 0 and gives a new deadline, but it does not roll back the token. If A's token=1 is still alive and you use token=1 again after the redrive, an old completion message could work on the new task. A fresh start of the count budget and reusing the work ownership generation are different matters.
What it looks like in the field
Even with the name DLQ, if no operator actually looks at it, you have only moved the failure to another place. You have to observe the quarantine count, the ratio by reason, the oldest quarantined work, the last investigation time, and the redrive results. In alerts, rather than putting in the whole work body, use the minimum identifier you need. If many work items were quarantined by a temporary outage, instead of reverting them all at once and creating load again, you must also limit the recovery speed.
AWS's DLQ guidance explains the flow of separating messages that could not be processed, analyzing the cause, and redriving them. It also warns about the impact of quarantine breaking the order for work that needs exact ordering. The dead row in this lab is a state of the local jobs, not a separate AWS queue, and its retention expiry and service permission policy also differ. You do not say that the product's features and guarantees are provided just because the word is the same.
In the comprehensive test, dispatcher A sends a real HTTP request, the receiving server commits the ID and the stock effect, and then A ends with exit code 73. The finally cleanup and finish are not run. Another process B reclaims the work at the expiry boundary with token=2 and sends the same ID again. The HTTP request really arrives twice, but the receiving stock increases by only 7 and the queue becomes done. The late completion request with token=1 must be False.
Do not expand this result to exactly-once processing of an external payment. The effect at that server happens once because the receiving server provided an implementation that atomically records the same ID. If an external provider does not support the same contract, you may need lookups, reconciliation, and manual confirmation. That you built the queue well does not atomically bind the external system's state as well.
The restart also has a clear scope. This verification is a process termination and resumption on the same local disk. It does not implement a backup to recover when the disk disappears or the DB is corrupted, consensus among several hosts, lease renewal, send cancellation, or a global request rate limit. It does not bring over the WAL setting from the earlier snapshot lab and instead builds a separate DB with a rollback journal suited to short queue writes. It is practice in choosing the storage mode according to the read requirements and write requirements.
What you will do in the next lab
In the 8-step comprehensive lab, you complete the retry function, persistent intake, claim, completion, redrive, and a single run. You check the claim race between two real dispatchers and termination after an HTTP receive, and you see whether wrong answers such as the expiry boundary, token reset, and a loose ACK are rejected, not just the right answer. The session is estimated at 110 minutes, so extend the time midway and keep any code you need separately.
Reference: SQS DLQ design, SQLite errors and rollback. The redrive in this lab assumes it is a local action that a person has already approved.