TT Lab
Get started
Learn Learning paths Courses

The snack machine died before ACK

Restart the Dispatcher That Died After ACK

Continue in TT Lab

In one line

Mark an unfinished item complete only after receiving a verifiable ACK, and when it fails, resend it with the same business ID while preserving the items that already finished.

Why this was needed

Even if you made the intake DB and the warehouse DB each safe, a single line in the delivery loop can ruin recovery. If you save sent=1 before sending the request, an item that died mid-delivery disappears from the next query. If you mark it after sending, an exit between the warehouse applying it and the marker causes a resend. Neither makes the send happen once. This lab chooses to allow resends and protect the receiving effect.

An operator should ask "up to which state is there evidence?" rather than "did this request succeed?" Being in orders, being in outbox as unfinished, being in the warehouse inbox, the ACK coming back to the publisher, and sent being saved are all different observations. One step does not automatically prove the next. If you organize this difference into a table during incident response, you reduce the mistake of deleting data to erase the uncertainty.

How it works

Process a small batch in order

pending reads the items with sent=0 by the outbox insertion-order seq. It does not sort by the ID's alphabetical order or by the second-level time. ID z can be accepted before a, and several items can have the same time. It returns only 1–16 at a time, and the query itself does not change any state. If you change them to complete as soon as you query them, you lose the items that have not been sent yet.

dispatch calls send for each item of the selected batch. When it receives an exact True ACK, it calls mark_sent. At the first failure it stops and propagates the original error. The items that finished earlier stay complete, and the failed item and the later items not yet attempted remain unfinished. If you run it again, it continues from that point. This lab deals with the order of a single delivery loop and does not implement claim, lease, and fencing, in which several workers take the same batch at the same time.

For example, among A, B, and C, suppose the completion marker was saved after A's ACK and B's response vanished; then B and C remain in the next batch. Even if B was already applied at the warehouse, it is resent with the same ID, so the inbox prevents the duplicate effect. C has not been sent yet. A policy that skips B's failure and completes from C onward is possible in some systems, but for work where order matters it means something different. This lab's stop-on-first-failure is an explicitly chosen contract.

Prevent unnecessary locks and argument mutation

Finish the pending query and close the SQLite write transaction before sending. This is so that, while send is slow or gets no response, you do not lock out intake of new orders as well. In the test, you check whether a separate DB connection can start BEGIN IMMEDIATE inside the send callback. This looks at whether a competing connection can actually enter the write boundary, rather than merely reading the value of con.in_transaction.

You also have to think about the fact that the other side can change the dict you passed to the callback. If the item you selected was B but the callback changes the argument's id to C, and you use that changed value as is for the completion marker, you can lose C. The teaching event has only strings and integers, so a dict copy is enough. You leave the marker with the originally selected value and send with the copy. If it is a different contract that includes nested objects, a shallow copy alone may not protect you.

Force a real failure order

The final check puts snack-7 into a temporary source DB and runs dispatch in a separate process. The HTTP server commits the receive ID and stock to a separate sink DB. Right after confirming the success ACK, the publisher exits with os._exit(73) in an after-delivery hook. finally and a normal close are not called. You check the exit code to see that the first process really exited at that point.

The parent checker reads from an independent connection whether an unfinished item remains in source and stock 7 remains in sink. It then starts a new publishing process. The requests the server observed must be two with the same ID and the final stock must be 7. In the second process, the outbox must change to complete. This is not a test that only recreates the state of a Python object; it is a test in which different processes reopen the same file.

In another case, the server closes the connection without sending the HTTP response after the receive commit. In this case the publisher must get an error and leave the item unfinished. When it connects again and sends the same item, the server ACKs it as a valid duplicate and the stock stays as it is. An ACK with the wrong ID, an oversized response, and a redirect are also not accepted as evidence of completion. send_http allows only this lab's loopback /events, does not use proxy environment variables, and does not follow redirects.

What it looks like in the field

In real operation, you do not judge the state from the count of successful sends alone. You look separately at the number of unfinished items, the age of the oldest unfinished one, the number of attempts, the number of receive duplicates, and the number of conflicts. If one item at the front is a permanent format error, the items behind it can stay blocked. Instead of infinite retries, you need to prepare alert, quarantine, fix, and reconciliation procedures. Quarantining is not a success, and you must also record the effect of changing the order.

Retry waits need a cap and jitter, but this dispatch runs only the finite batch of a single call. Do not describe it as having implemented an automatic retry scheduler or high-availability work distribution. The source and sink files also disappear when the LabHub session ends. What was verified here is the case of reusing the same disk after a process exit. Host power cutoff, disk failure, backup recovery, and performance over the internet are separate verifications.

What you will do in the next lab

In worker.py, you implement in order event validation, store initialization, atomic intake, the unfinished query, duplicate receipt handling, the completion marker, finite batch delivery, and HTTP ACK verification. The provided checker creates the temporary DB and loopback server itself and runs your code. An empty function skeleton is valid only syntactically and does not stand in for an answer. In the quiz after the lab, you sort out the scope that the result of two deliveries and one application demonstrates and the operational responsibilities that remain.

Further reading in the official docs