TT Lab
Get started
Learn Learning paths Courses

The snack machine died before ACK

Two Deliveries, One Snack Order

Continue in TT Lab

In one line

You do not refuse to receive duplicate messages; you commit the receive record together with the work so that the effect of the same business ID is not applied again.

Why this was needed

The warehouse prepared 7 snacks and sent the ACK, but the receiver's connection dropped. The receiver is left with two hypotheses. The warehouse may have done nothing, or it may have already done it and only the response vanished. Giving a longer timeout does not make this window disappear completely. Sooner or later a connection drops and a process restarts. If you treat a certain failure and a failure whose outcome is unknown as the same thing, you end up with duplicate shipments or omissions.

For the receiver to send an uncertain item again, the warehouse has to be able to recognize the same work. If you put the ID in an in-memory set, it can prevent duplicates while running, but the next process has no memory of it. If you write only the ID to a file first and then update the stock, the stock is lost if it exits midway, and if you write the stock first and then record the ID, the stock can increase twice. The receiving side also needs a single atomic boundary.

How it works

The inbox is saved together with the work

The warehouse's inbox holds id and qty. The single row of stock has the cumulative quantity total. receive opens a write transaction with BEGIN IMMEDIATE and looks for a record with the same ID. If it is an ID it sees for the first time, it inserts into the inbox, increases total, and commits. If it is the same ID and the same quantity, the work is already applied, so it finishes without increasing again. If it is the same ID but a different quantity, it rejects with Conflict and preserves the earlier work.

Here, a return value of True from receive means that this call applied a new effect. False means it confirmed a valid duplicate, not a failure. The HTTP receiver returns an accepted=true ACK with the same ID in both cases. If you keep answering a duplicate request with a failure because there was no new effect, the publisher retries forever. It is important not to mix the internal function's "did it apply anew?" with the communication response's "did it accept this work?"

처음 snack-7 / 7 → inbox 1행, total=7, 새 적용 True, ACK 성공
다시 snack-7 / 7 → inbox 1행, total=7, 새 적용 False, ACK 성공
다시 snack-7 / 8 → 내용 충돌, 기존 total=7 보존, 성공 ACK 없음
다른 snack-8 / 7 → inbox 2행, total=14, 별개 업무 ACK 성공

This table shows that we do not merge different work just because the content is the same. Even if you attach a content fingerprint such as SHA-256, the responsibility for defining the same work does not go away. A fingerprint is a means of comparing whether the content changed under the same key; it is not a business rule that decides whether a customer ordered twice. In this small lab the content is a single integer quantity, so we compare it directly. When you widen it to structured content, you also have to consider normalization and schema versions.

When may you produce the ACK

You produce the ACK after receive's commit has finished. You do not treat HTTP status 200 alone as confirmation of the work the publisher requested. You check that the id in the response JSON equals the requested id and that accepted is exactly the boolean True. You do not let the string "true" or the number 1 pass as truthy. This is because you cannot use another order's success response as evidence for this order. In a real service, you would also separately check that the peer is authenticated and the protocol version of the response.

This server is for teaching and runs only on an ephemeral port of 127.0.0.1. It has no authentication or TLS and is not an API for external exposure. We used HTTP to observe real requests, responses, and connection closes rather than just mock errors between function calls. Do not over-interpret two independent DBs on the same host as two services across a network. Public network latency, clock skew between servers, and device loss are not included in this test.

What if the same ID arrives concurrently

Another reason the duplicate check and the write must be in one transaction is contention. If two requests both read "not there yet", each runs the work, and then later try to save only the ID, a gap opens between the check and the execution. SQLite's BEGIN IMMEDIATE is used to serialize write contention. If another write is already in progress, it may wait or raise a lock error. Do not turn that into "this work is complete".

A unique key constraint helps prevent duplicate rows, but it does not undo a separate external side effect. The property of this lab, changing stock in the same transaction as the inbox insert, cannot be attached as is to an external payment call. External work needs a separate design such as the other side's idempotency contract, compensation, and status lookup. Also, if you do not decide how long to keep receive IDs, an old resend can be processed as new work.

What it looks like in the field

Notifications, customer data integrations, and real-time event consumers often fail on the assumption that "resends are rare". When a process is replaced right after a deployment or a proxy cuts a response, resends occur even for normal work. If you observe the number of communication requests and the number of actual business applications separately, you can tell these apart. 2 requests and 1 piece of work can be the normal result of deduplication working. 1 request and 0 pieces of work cannot simply be hidden behind a low average latency metric.

Leaving the business ID in the error log helps, but unconditionally leaving customer names, addresses, and payment information creates a different risk. This lab uses only fictional IDs and quantities. In a production design, you should keep only the correlation you need and decide the retention period and access permissions. Do not substitute dumping all resent content to a log for an idempotency store.

What you will do in the next check

In the quiz right after this, you distinguish a valid duplicate, a content conflict, separate work, and the meaning of the ACK. In the next module, you learn where the publisher's completion marker must live and the resume order for a failed batch. In the lab, you end the process after the inbox insert, after the total update, and after the commit, and read the state from an independent connection each time.

Further reading in the official docs