Idempotency — Two Clicks, One Charge
Resume imports from the last committed checkpoint: design principles
In one line
Combine the source fingerprint and the checkpoint to prevent wrongly resuming with a different file.
Why this was needed
A job importing thousands of rows died partway. The operator swapped the file and ran it again with the same job id, and the result had the first part from the old file and the rest from the new file. If you store only the number of the processed row, you cannot confirm the identity of the input. You must keep the fingerprint of the source bytes together with the last committed position.
How it works
The input is a JSON array with no duplicate ids, and each row has an id and a value. The SHA-256 of the source bytes and the checkpoint are stored in imports. The insertion of each item in a batch and the increase of next_index are in the same transaction. If an exception occurs in a middle row, that whole batch is rolled back and the batches that finished earlier remain. A rerun starts from the stored index, but it is rejected if the source fingerprint differs.
원본 bytes → 지문 확인 → next_index → 배치 INSERT + 체크포인트 COMMIT
└ 중간 실패 → 이번 배치만 ROLLBACK
Worksheet: read the contract and predict the failure
What follows is not an answer sheet for memorizing the implementation, but a step-by-step code review. Each changed fragment deliberately breaks the contract. Note that normal cases may still pass after the change. Before running, predict which input, exception, or state you would have to observe to reveal the difference, and after implementing, compare that prediction with the result.
1. Check the contract of the input rows
parse_rows(raw) reads JSON array bytes. Each item has an id (a non-empty str) and a value (an int, not bool), and duplicate ids are forbidden. A violation is a ValueError. It returns a list of rows that have only {id,value}.
Basis for the judgment: if a duplicate id is silently overwritten by the last row, the result of the import cannot be predicted.
The wrong changed fragment to review:
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
2. Fix the fingerprint of the source bytes
source_digest(raw) is the hex string obtained by applying SHA-256 to the bytes. It does not normalize the JSON.
Basis for the judgment: the contract allows resuming only for the same source.
The wrong changed fragment to review:
hashlib.sha256(raw.replace(b" ",b""))
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
3. Store the input and the checkpoint
init_db(path) idempotently creates imports(id TEXT PRIMARY KEY,digest TEXT NOT NULL,next_index INTEGER NOT NULL) and items(batch TEXT NOT NULL,id TEXT NOT NULL,value INTEGER NOT NULL,PRIMARY KEY(batch,id)).
Basis for the judgment: store the same row id from different import jobs separately.
The wrong changed fragment to review:
CREATE TABLE imports
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
4. Keep it from resuming with a different source
begin(path,batch,digest) stores next_index=0 and returns 0 for a new job, returns next_index for an existing job with the same fingerprint, and raises ValueError for a different fingerprint.
Basis for the judgment: even with the same job id and row number, the input file may be different.
The wrong changed fragment to review:
if False:
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
5. Commit the batch and the position atomically
apply_chunk(path,batch,start,rows,fault=lambda index:None) runs only when the current next_index==start, and otherwise raises ValueError. It puts rows into items in order and calls fault(overall index) after each insertion. If everything succeeds, it stores next_index=start+len(rows) and returns it.
Basis for the judgment: if you commit separately for every row, the checkpoint and the row state get out of step.
The wrong changed fragment to review:
db.commit()
fault(start+offset)
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
6. Look up the current position
checkpoint(path,batch) returns next_index, and None for a job that does not exist.
Basis for the judgment: read the last committed position, not the position of the last attempt to process.
The wrong changed fragment to review:
return 0 if row else None
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
7. Separate the results per job
values(path,batch) returns the (id,value) tuples of that batch in ascending id order.
Basis for the judgment: make batch a condition so that rows with the same id from another job do not get mixed into the result.
The wrong changed fragment to review:
WHERE batch!=? ORDER BY id
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
8. Continue safely after a failure in the middle
import_all(path,batch,raw,size=2,fault=lambda index:None) validates that size is a positive int (not bool) and uses parse_rows, source_digest, and begin. It processes the remaining rows size at a time with apply_chunk and returns the final checkpoint.
Basis for the judgment: check the scenario of resuming after the second batch fails while preserving the success of the first batch.
The wrong changed fragment to review:
begin(path,batch,source_digest(raw))
index=0
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
What it looks like in the field
This is a lab for small data that reads the whole input into memory. It does not exaggerate into an engine that parses large files in a streaming way. It is a conservative contract in which even a whitespace difference in the source changes the byte fingerprint and resuming is rejected. Side effects of external APIs do not go into this DB transaction.
What you will do in the next lab
The eight steps connect into one runnable deliverable. Check the contract of the input rows → fix the fingerprint of the source bytes → store the input and the checkpoint → keep it from resuming with a different source → commit the batch and the position atomically → look up the current position → separate the results per job → continue safely after a failure in the middle.
Each step checks not the fact that a function or file exists but the actual return values, exceptions, and state changes. After you see the answer, deliberately change a boundary comparison or the cleanup code and check which test fails. Explain why the earlier tests are kept in the next step too, and write down one operational condition that this lab does not guarantee.