TT Lab
Get started
Learn Learning paths Courses

The snack machine died before ACK

An Empty Log Does Not Mean You Are Up to Date

Continue in TT Lab

In one line

The cursor is the position to receive again, and the retention boundary is the evidence for judging whether that position still exists. Getting an empty response does not mean you are up to date.

Why this was needed

Suppose you show the stock of a club snack warehouse on a display board. Receipts are recorded as positive deltas and distributions as negative ones. The warehouse server kept running, but the display board was off over the weekend. On Monday, when it reconnects by sending its last processed number, it expects to receive the missed events as usual. But the operator deleted the early part of the weekend log to save disk. If the server sends only the remaining events, the display board produces plausible numbers that permanently drift from the real stock.

This is not a connection failure. Even if the TCP connection opens and the HTTP response succeeds, the fact that the past is gone does not change. The outbox in the previous module preserved the intent that was still to be delivered. This module is about recovering the current state even after old delivery records were deliberately thrown away. Keeping logs permanently and unconditionally does not avoid the retention cost, and resetting any state to initial values loses real work. That is why you design fast replay and full state recovery as separate paths.

How it works

In the lab, stream numbers increase from 0, and the last number when there are no events yet is expressed as -1. last is the last number applied, and floor is the earliest number the server still retains. The events that normally remain are consecutive from floor to last, and if all were truncated, floor is last+1. This is this lab's integer contract. Do not substitute it as is for another product's time-based IDs or partition offsets.

Server state Client last Judgment
floor=4, last=8 3 Replay from 4 is possible
floor=4, last=8 2 3 is missing, so a snapshot is needed
floor=9, last=8 8 Already up to date, an empty batch is normal
floor=9, last=8 -1 The whole log was truncated and recovery is needed

The key inequality is whether the client's last is less than floor-1. If it is equal, the next number is floor, so it is safe. If it is less, at least one event is missing. A cursor ahead of the server is also not accepted as normal. It may be that you connected to the wrong store or that the local state is from a different generation, so you must investigate it as a separate error. A single inequality sign can make data duplicate or go missing, so you test each boundary value.

The source store uses three tables: checkpoint, stock, and events. events is the individual changes to replay, stock is the cumulative stock, and checkpoint holds the generation, the last number, and the retention boundary. trim(through) deletes only the events at or below through and changes floor to through+1. It does not change the stock or last. Reducing the replay history and canceling business that already happened are different. Even when events drops to 0, the single checkpoint row is kept, so you can distinguish a warehouse that was empty from the start from a warehouse whose past was cut off.

When you query, you also look at the metadata and the events within a single read transaction. A structure that first judges the retention range to be sufficient, then lets another write cut the log, and only then reads the events has a different check time and use time. Based on SQLite's explanation of isolation between separate connections, you keep the same read point. If you do not actually use the isolation the library provides, the reason for storing the tables separately disappears.

What it looks like in the field

If the display board is off four hours every day and the replay log keeps only two hours, then however much you polish the reconnect implementation, you will often need a full recovery. Here you have to think together about the maximum offline time the product allows, the event volume, and the time to create and transfer a snapshot. Increasing the retention reduces the frequency of recovery but can raise the storage cost and the scope of retained personal data. Conversely, for a service with a small snapshot, a short log and a fast full recovery may be simpler.

In the diagnostic log, do not just write that the connection succeeded; leave the request generation, the client last, the server floor and last, and the recovery path you chose. This lab uses only integer stock with no customer personal information. In a real service, instead of printing the whole state to a log, you should minimize the identification scope and the diagnostic metadata. A snapshot that contains business data is not a file to put at just any public link.

A single replay returns at most 16. This is not a number that guarantees the full recovery is complete; it is the resource budget of one call. If 20 remain, it is normal for four more to remain after the first call. You must distinguish an empty batch from a missing error so that the caller can choose its next action. If you mark a single success as fully complete while a lot remains, the progress indicator deceives the user.

What you will do in the next check

This module first judges the boundaries with the quiz, and in the comprehensive lab that follows you implement append, trim, and replay. You compare whether an error occurs when you send an old cursor against a fully truncated log and whether the latest cursor gets an empty batch. You also check that after trimming, the new numbers do not go back to 0. In the next reading, you bind the stock and the cursor used for recovery to the same point in time.

Reference: SQLite isolation between connections. The meaning of floor and last, the batch size, and the error classification are the application protocol that the warehouse lab above defined.