The snack machine died before ACK
Do Not Read a New Semester with Yesterday's Cursor
In one line
Even with the same number, it may be an event from a different epoch. Recovery is the process of installing a snapshot of an allowed epoch and then following, for a finite number of steps, from the next number.
Why this was needed
On the first day of the school festival, the stock records went up to number 800. On the second day they opened the warehouse fresh and started the log numbers from 0. If you send only the number 800 left on the display board, the server may mistake it as having already processed even the events that are yet to happen. Or a replica may discard the new log as a duplicate of the past. This problem arises because nobody stated the range in which the number is unique.
Here we use epoch as a generation identifier such as day-A or day-B. Within one epoch, numbers are not reused, and a different epoch is assigned only when a new epoch is explicitly started. The strings in the lab are examples to make the explanation easy, and they do not mean that anything is safe as long as the date is the same. A real system needs an identification policy that distinguishes restores, regenerations, and operating environments. If you let the epoch be chosen from arbitrary user input, it could even become a security problem of fetching another customer's state.
How it works
install_snapshot takes expected_epoch in addition to snapshot. You can install only if snapshot.epoch equals expected_epoch. expected_epoch is a value passed in from an already trusted connection setting, and if you pull it straight out of the JSON you are checking, the check becomes circular. That it is a well-formed string is also not evidence that you have permission to read that warehouse. Authentication, authorization, and transport protection must be guaranteed separately by the real service.
An older snapshot of the same epoch is rejected as a Conflict. If last is the same but total differs, that is also a conflict. If the epoch, last, and total are all the same, there is no need to reinstall, so it returns False and does not delete the remaining individual events either. If it is a snapshot of another allowed epoch, it can replace even with a smaller number. You must not reject a new epoch merely on the numeric comparison that day-B's 0 is smaller than day-A's 800.
Once you install the snapshot, the individual events before it no longer remain in the replica. For example, a snapshot with total=13, last=5 tells you only the sum from 0 to 5 and does not tell you what the delta of number 4 was. If (4,2) comes in late, you cannot prove it is a duplicate with the same content, so you classify it as CoveredBySnapshot. If you quietly ignore every past number, you can hide a content conflict. Conversely, if both the number and the delta of an individual event you kept are the same, it is a verifiable duplicate and the effect is not added again.
A new batch must have consecutive numbers. If the last applied is 5, the first new number is 6. If you receive a batch with only 6 and 8 and raise the cursor to 8, 7 disappears. Also, if you commit the effect of 6 and then raise an error at 8, confusion arises when the caller retries thinking the whole batch failed. This lab applies a whole batch of up to 16 as a single transaction and first checks the format and the internal order. It binds the processing record, the stock, and the last number into the same boundary.
What it looks like in the field
sync_once is a single synchronization attempt. It checks that the server epoch is the expected value and requests a replay with the replica's cursor. If it is out of range or the epoch differs, it receives a snapshot, installs it, and then reads only events greater than snapshot.last. It returns the state after applying up to 16, and it does not promise that it has fully reached the latest server state. The caller compares the resulting cursor with the server cursor at the observation time and decides whether to keep following or to show the user the delay.
In particular, the source log can be cut again while you install the snapshot. Suppose the snapshot initially held last=20, but after the install the source wrote 21 and deleted it as well; then you cannot replay 21. In this case this sync_once propagates ResyncRequired again. It does not erase the valid state of last=20 that was already installed, and it lets the next finite call receive a newer snapshot. If you hide the failure or retry endlessly inside a catch, you only burn resources without catching up to the source's fast retention cuts.
sync_files is responsible for the boundary of opening and closing two different local files. Even if the path strings differ, they can be the same file through a hard link, so you also check the real file identity. If the source does not exist, it does not create an empty DB and proceed as if the recovery had succeeded. Even if opening the replica fails, it closes the source connection it already opened. The file paths assume a disposable folder controlled by the learner, and it does not implement defense against a path race on a production server where an attacker swaps links at the same time.
What you will do in the next lab
In the 8-step comprehensive lab, you complete apply_batch, sync_once, and sync_files. At the end, you end a real child process with exit code 73 and reopen the DB with an independent connection. At each point before and after the commit, you compare whether the cursor, stock, and individual events are entirely the earlier state or entirely the new state. It verifies restarting with the same file but does not verify disk loss, consensus between servers, automatic failover, TLS, or user permissions. When you extend to the next assignment as well, you need work that widens the boundary you verified.
Reference: Python sqlite3 connections and transactions, Python os.path.samefile. The epoch, the conflict classification, and the finite recovery loop are contracts of this learning protocol.