Idempotency — Two Clicks, One Charge
Deleting an idempotency key ends its guarantee: design principles
In one line
You distinguish completed records from in-progress records, and implement a retention period and batch cleanup.
Why this was needed
When the table grew large, all the old idempotency keys were deleted. The records of requests that were still in the middle of a payment disappeared too, and the retries came in as new requests. The retention period of a completed response and the ownership of an in-progress job are not the same expiry policy. Cleanup is not a simple DELETE; it is a state transition that changes the scope of the guarantee.
How it works
A key is reserved as pending and changed to done only with the same body fingerprint. A completed response is replayed until the expires time. Cleanup deletes only rows that are done and have expires at or before now, limited by id order and limit. Pending is never deleted by this cleanup function, no matter how old it is. The last lab has you reserve the same key before and after cleanup, and directly confirm the limitation that once a record is deleted, the key is treated as a new request.
pending → done → expires 도달 → 제한된 purge → 같은 키가 새 요청이 됨
pending ─────────────────→ purge 대상 아님
Worksheet: read the contract and predict the failure
What follows is not an answer sheet for memorizing the implementation, but a step-by-step code review. Each changed fragment deliberately breaks the contract. Note that normal cases may still pass after the change. Before running, predict which input, exception, or state you would have to observe to reveal the difference, and after implementing, compare that prediction with the result.
1. Separate completion from expiry
init_db(path) idempotently creates keys(id TEXT PRIMARY KEY,fingerprint TEXT NOT NULL,status TEXT NOT NULL,response TEXT,expires REAL NOT NULL).
Basis for the judgment: do not judge that the work is finished just by looking at the expiry time.
The wrong changed fragment to review:
CREATE TABLE keys
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
2. Reserve an in-progress key
reserve(path,key,digest,expires) inserts a key that does not exist as pending,response=NULL and returns True. An existing key is not changed, regardless of expiry, and it returns False.
Basis for the judgment: control record deletion and key reuse separately and explicitly.
The wrong changed fragment to review:
INSERT OR REPLACE INTO keys
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
3. Complete only the current request
finish(path,key,digest,response) stores response as JSON and changes it to done, returning True, only for a row that has a matching key and fingerprint and is pending; otherwise it returns False.
Basis for the judgment: compare even the fingerprint so that a worker with a different body cannot overwrite a completed response.
The wrong changed fragment to review:
AND status IN ('pending','done')
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
4. Replay only a valid completed response
fetch(path,key,now) parses the response of a row that is done and has expires>now as JSON and returns it. Otherwise it returns None.
Basis for the judgment: pin down the contract of no longer replaying at the boundary now==expires.
The wrong changed fragment to review:
AND expires>=?
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
5. Limit the cleanup candidates
expired(path,now,limit=10) checks that limit is an int from 1 to 100 (not bool). It returns up to limit ids that are done and have expires<=now, in ascending id order.
Basis for the judgment: if you delete the whole table at once, the lock time gets long and it is easy to mix in in-progress rows.
The wrong changed fragment to review:
return [r[0] for r in db.execute("SELECT id FROM keys WHERE status='done' AND expires<=? ORDER BY id DESC
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
6. Put selection and deletion in the same transaction
purge(path,now,limit=10) picks candidates with the same limit validation, condition, and ordering as expired, deletes them in one transaction, and returns the list of deleted ids. It does not delete pending.
Basis for the judgment: do not create a gap in which the state changes between looking up the candidates and deleting them.
The wrong changed fragment to review:
ids=[r[0] for r in db.execute("SELECT id FROM keys WHERE expires<=?
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
7. Actually count by state
counts(path) is {pending: count, done: count}. Even when there is no row in a state, the key must exist with 0.
Basis for the judgment: you must be able to observe that unfinished rows did not disappear after cleanup.
The wrong changed fragment to review:
"pending":0
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
8. Confirm that the scope of the guarantee ends after deletion
retention_cycle(path,key) reserves with expires=10,digest='v1' and completes with {receipt:1}. After purge at now=10, it reserves the same key anew with digest='v2',expires=20 and returns the resulting bool. The function assumes an empty DB.
Basis for the judgment: do not hide the limitation that after cleanup the key becomes a new request, and reproduce it directly.
The wrong changed fragment to review:
purge(path,9)
Compare it with the public contract of the function that contains this fragment. If a single success case cannot tell the difference, choose as the observation target an input that should be rejected or the state left after a failure.
What it looks like in the field
It does not guarantee an exactly-once effect for an unlimited period. An external system may resend an old request, so the retention period must be matched to the provider's retry policy. The lease and compensation policy that recovers abandoned pending rows is the responsibility of a separate lab, and this cleanup function does not guess and delete them.
What you will do in the next lab
The eight steps connect into one runnable deliverable. Separate completion from expiry → reserve an in-progress key → complete only the current request → replay only a valid completed response → limit the cleanup candidates → put selection and deletion in the same transaction → actually count by state → confirm that the scope of the guarantee ends after deletion.
Each step checks not the fact that a function or file exists but the actual return values, exceptions, and state changes. After you see the answer, deliberately change a boundary comparison or the cleanup code and check which test fails. Explain why the earlier tests are kept in the next step too, and write down one operational condition that this lab does not guarantee.