TT Lab
Get started
Learn Learning paths Courses

The snack machine died before ACK

Connection recovery is not business recovery

Continue in TT Lab

In one line

A dropped connection can be reopened, but to decide whether to redo work that was already done, you must save the business result and the processing cursor in the same transaction.

Why this was needed

A snack vending machine on a space station receives events. EVENT 0 7 means increase the stock by 7. The power went out just after the receiver saved the stock, so it could not send the ACK. To the publisher, "processed but the reply vanished" and "not processed at all" look exactly the same. A successful TCP reconnect alone cannot tell the two situations apart. You need a design that resends an uncertain event while the receiving side recognizes an event it has already processed.

The Cursor in the previous course is the memory of a running object. When you start the process anew, that memory disappears too. This time you save the sequence number and the total in a SQLite file, and you terminate the real receiving process right before the ACK. You then check whether a second process opens the same file and picks up where the first left off. This is not a lab for building a payment server; it is an experiment that exposes the recovery boundary through the ordered stock changes of a single stream.

How it works

Turn the bytes into a contract first

The practice protocol is ASCII EVENT, a space, the sequence number, a space, the delta, and a single LF. For example, EVENT 0 7 followed by a newline is one event. The sequence number is a consecutive integer starting at 0, and the delta ranges from -1000 to 1000. Leading zeros, a + sign, -0, CRLF, and a final fragment with no newline are not allowed. The strict grammar was set by the author so that you distinguish "it converted to a number, so it's fine" from "it is the message we agreed on". If another protocol allows CRLF, you must follow that contract.

This message is not WebSocket or SSE. Those two standards each have their own separate framing and reconnection rules. Here we lay a small line protocol on top of a TCP byte stream to focus on processing semantics. The line limit is 64 bytes, and the StreamReader limit of open_connection is also set to 64. If you keep going from the next line after reading a bad message, you can hide a command that was actually missed, so you terminate the connection. HTTP authentication, TLS, and per-user permissions are not implemented, so you must not expose this to the internet as it is.

Keep the business effect and the cursor in one place

The single row with id=1 in the checkpoint table stores last and total. When nothing has been processed yet, last=-1 and total=0. The ledger uses seq as its primary key and keeps the delta that was processed. This lab does not delete the ledger; a long-term retention and compaction policy for the processing history has to be designed separately.

apply_event opens a write transaction with BEGIN IMMEDIATE and reads the current last. If seq is last+1, it records it in the ledger, changes total, advances last, and then runs COMMIT. On an error midway it does ROLLBACK and returns the original error to the caller. The lab's fault hook runs after total is changed and before last is changed. At this point the checker observes the partial change from the same connection, injects the error, and then checks from another connection that no change was left behind. It also terminates a separate process immediately at this hook to check the case where finally does not run. A mere printout saying "the exception was caught" does not stand in for atomicity.

If the same seq and delta arrive again, it does not add the result and returns False. If the seq is the same but the delta differs, it is a Conflict. This is so that a publisher that reuses an event ID while changing its content is not quietly accepted as a duplicate. If seq is greater than last+1, it is a Gap. If you processed up to 4, received 6, and jumped to last=6, you would later mistake the 5 that arrives for already processed. You must stop for now and recover the missing range.

Here we turn off the automatic BEGIN with isolation_level=None of sqlite3.connect and state BEGIN, COMMIT, and ROLLBACK explicitly in SQL. Do not mix this with the autocommit=False mode of Python 3.12 and create a nested BEGIN. with con is not syntax that closes the connection itself, so the owner closes it in finally. open_store does not overwrite existing rows with initial values. If you reset to last=-1 on restart, writing the file would have been pointless.

The ACK comes after the commit, and the resume continues after the cursor

consume_line returns the ACK sequence number and LF after parsing and apply_event have succeeded. It ACKs a duplicate with the same content too. This lets the publisher finish its retries without adding the result again. If you send the ACK before committing and then die, the publisher may erase the event. Conversely, when the ACK is lost after the commit, a duplicate may occur, but it can be filtered out using the processing history.

run_client opens the store, connects over TCP, and sends RESUME last and LF. Reconnecting is done explicitly through a new run_client call. We do not hide an infinite automatic retry loop. In the final check, the first receiver saves 0 and 1 and then exits with os._exit before ACK 1. The second process sends RESUME 1 and receives the 1 that the server deliberately resends and the new event 2. You compare together the requests and ACKs observed at the server and the ledger and total read over an independent connection. This is not a mock restart that swaps out a running object.

Outside the retained log, recovery works differently

replay returns, from the publisher's finite log, up to limit items greater than the cursor. If the current retention start is 10, last=9 can read from 10, but last=8 is missing the needed 9, so it is ResyncRequired. If you hand over only the latest items and call it a success, that is a silent loss. A real product must either resynchronize from a consistent snapshot and that snapshot's cursor or provide an explicit recovery failure. This lab does not implement creating or installing the snapshot itself.

An empty log, without retention information, cannot distinguish "it was empty from the start" from "everything was truncated", so it is a ValueError in this API. The input is limited to 1–128 consecutive items and the batch to 1–16. A real log service needs separate metadata such as a lower and upper retention bound. Redis XREAD also reads items after a specified ID, but you must not interpret this lab's integer sequence numbers and exceptions as Redis's actual behavior.

What it looks like in the field

Notifications, collaboration screens, and AI streaming get dropped by deployments and network changes even when the connection is kept open for a long time. Reconnection, replay, the processing cursor, and the retention policy must be designed together. As the number of connections grows, retries also need exponential backoff, jitter, a maximum count, and per-account resource limits. This time we test only two connections and a small log, so do not use it as evidence of internet latency or large-scale throughput.

The atomicity of the total and the cursor in the same SQLite transaction does not extend to an external mail or payment API. A new failure window also appears between the DB commit and the HTTP request. For such systems you have to learn, separately, designs such as idempotency keys on the receiving side and an outbox, as well as reconciliation procedures. The explanation "there is an ACK, so it is exactly once across the whole distributed system" is wrong.

The lab files survive a process restart but disappear when the LabHub session ends. This process-failure check does not test host power loss, disk failure, or backup recovery. SQLite's durability also rests on assumptions about the filesystem and sync behavior. Report the failure types you observed and the types you did not guarantee separately.

What you will do in the next lab

You complete client.py in this order: frame parsing, then the store, then atomic processing, then judging the retention range, then the ACK after the commit, then connection reclamation, then a real process reconnect. timeout is the wait limit of each network await, not a total deadline that includes the whole session and the DB work. The SQLite calls are synchronous, and this small experiment has a single receiver. You should avoid a design that runs a long DB operation as is on a production event loop.

Further reading in the official docs