The Same Transfer Went Out Twice — Build the Side That Stops It
Goal
You build yourself a layer that makes the money go out only once even if the same transfer request comes in twice because of client retries. You attach the idempotency key table and UNIQUE constraint, the normalized fingerprint of the request body, the in-progress state, response replay, and the key scope and retention period, and replay a day's request log to prove that duplicate transfers are 0.
Why it matters
A client that did not get a response sends again. There is no way to stop this, and you must not stop it. RFC 9110 does not treat POST as idempotent, so the protocol gives no guarantee — the application builds the guarantee. The skeleton of how to build it is organized by an IETF draft (draft-ietf-httpapi-idempotency-key-header). The client makes the key, the server makes the fingerprint, a retry of a completed key replays the stored response, and a retry of an in-progress key is answered with a conflict. What is hard in this lab is not the code but the boundaries. If you look up and then insert, you miss concurrent retries; if you write the key as completed first, the money never goes out when it dies; and if you do not look at the fingerprint, a request whose amount changed is quietly ignored. The grader does not trust your wording. It sets up an account database that the grader itself built in a temporary directory, actually runs your script with different accounts, amounts, and keys each time, and compares the response JSON, the transfer table, and the balances against the values it measured itself.
Steps
- Create and run /root/idem/gen_requests.py to make /root/idem/idem.db. It holds 8 accounts, 120 requests (100 distinct keys), and that day's 120 transfers.
- Tally the duplicate transfers created by retries and write them to /root/idem/dup_report.json as requests, unique_keys, duplicate_keys, extra_transfers, double_paid, and keys.
- Create /root/idem/idem_api.py and use the UNIQUE constraint of the idempotency key table so that a second request with the same key cannot create a transfer. If the key is different, do not block it.
- Make idem_api.py store the normalized fingerprint of the request body and reject with 422 when the same key comes with a different body. The same body with only the field order different is the same request.
- Split the key record of idem_api.py into two steps, claim (in_progress) and completed, and make it answer a retry of an in-progress key with 409 in_progress. You must be able to create the situation of dying before processing with
--crash-after-claim. - Make a retry of a completed key replay the stored response as it is. The status is the same as the first response, replay is true, and the body must not differ by a single character.
- Widen the key scope to (client_id, endpoint, idem_key) and delete keys past the retention period with
--purge-before. Write the policy in /root/idem/policy.json. - Replay all 120 requests without dropping any using /root/idem/replay_day.py to produce /root/idem/day.db and /root/idem/result.json, and report in /root/idem/idem_report.md in four sections.
Notes
- Execution contract:
python3 /root/idem/idem_api.py --db <DB> --request <요청 JSON>(the placeholders are the DB file and the request JSON file) writes one chunk of response JSON to standard output and ends with exit code 0. If the request file cannot be read, it is 3. - Request JSON:
{"client_id": …, "endpoint": …, "idem_key": …, "body": {"src": …, "dst": …, "amount": 정수, "currency": "KRW"}}(in this and the next example, the Korean words mean integer, string, and object). - Response JSON:
{"status": 정수, "replay": true|false, "reason": 문자열|null, "body": 객체|null}. If it was newly executed, it is status 201 and replay false, and body holds a transfer_id. - Table names and columns:
account(acct_id, holder, balance)andtransfer(transfer_id, src, dst, amount, currency, created_at); the idempotency key table is for you to decide. The grader pre-fills only the account table and expects your script to create the rest. - Locking: Python's sqlite3 uses deferred transactions by default. For places where writing is certain, such as claiming, use
BEGIN IMMEDIATE, and handle waiting withPRAGMA busy_timeout. - Common mistakes: inserting after a lookup (you miss concurrent retries), writing the key as completed first, computing the fingerprint from the raw string, and recomputing the response when replaying.
- The 72-hour retention period and the three-column scope are assumptions of this lab. The IETF draft only says to publish the expiry policy in documentation and does not fix a number.
- Do not build a load test. The budget for one grading is 60 seconds and the Pod has 2 cores.
Build that day's request log
Create and run /root/idem/gen_requests.py to make /root/idem/idem.db. It has 8 accounts, 120 requests (100 distinct keys), and that day's 120 transfers.
First create /root/idem and build the sqlite DB with python3 in it. There are three tables: account, req_log, and transfer_v1. req_log is the requests the gateway received as they were, so retries are in it as one line each, and transfer_v1 is the transfers those requests actually created.
Count the duplicate transfers created by retries
Write requests, unique_keys, duplicate_keys, extra_transfers, double_paid, and keys to /root/idem/dup_report.json. duplicate_keys is the number of keys for which two or more transfers were created, and keys is the list of those keys.
extra_transfers is the count left after removing the first transfer for each key. double_paid is the sum of the amounts of those leftover transfers. If you number which transfer it is within each key with ROW_NUMBER() OVER (PARTITION BY … ORDER BY …), available from SQLite 3.25, you get it in one go.
Block the second request with the same key
Create /root/idem/idem_api.py and use the UNIQUE constraint of the idempotency key table so that a second request with the same key cannot create a transfer. If the key is different, do not block it even with the same body.
The approach of looking up and inserting if absent lets two requests that arrive at the same moment both pass. INSERT straight into a table with the key as primary key, and use the constraint violation exception (sqlite3.IntegrityError) as the signal for "it already exists." For the second response, you can set replay to true or answer with 409.
Reject it when the key is the same but the body is different
Make idem_api.py store the normalized fingerprint of the request body together with the key and reject with 422 when the same key comes with a different body. The same body with only the field order different must be treated as the same request.
If you compute the fingerprint from the raw string, a difference in only field order or whitespace makes it a different request. The core of the normalization RFC 8785 specifies is key sorting and whitespace removal. In Python, you can approximate it with sort_keys and separators of json.dumps, and feed those bytes to sha256.
In-progress keys and concurrent retries
Split the key record into two steps, claim (in_progress) and completed, and make it answer a retry of an in-progress key with 409 and the reason in_progress. --crash-after-claim must claim only and end with a nonzero code without a transfer.
If you write the key as completed first, then when it dies during the transfer, the money did not go out but the retry receives "already processed." If you separate the claim transaction and the execution transaction, the intermediate state remains in the table. The grader also sends two retries that arrive at the same moment — exactly one of the two must be newly executed.
Replay the stored response as it is
Make a retry of a completed key return the stored response as it is. The status is the same as the first response, replay is true, and the body must not differ from the first response by a single character.
Replay is not recomputing. Store the response body whole at completion time and take it out as it is. If you recompute, the balance and time change and the customer's screen looks different twice. An in-progress key must still be 409.
Decide the key scope and retention period
Widen the key scope to (client_id, endpoint, idem_key), and make --purge-before <RFC 3339 시각> (the placeholder is an RFC 3339 timestamp) delete keys older than that and output {"purged": 개수} (where the Korean word is the count). Write the policy in /root/idem/policy.json as scope, retention_hours (72), on_fingerprint_mismatch (422), and on_in_flight (409).
If the key is global, then when another customer happens to send the same string, they get someone else's response. If you change the primary key to three columns, a scope arises. If you delete the key after the retention period, the same request afterward is a new request — so the retention period must be longer than the client's retry limit.
Replay a full day to prove it
Replay all 120 requests without dropping any using /root/idem/replay_day.py to produce /root/idem/day.db and /root/idem/result.json, and write /root/idem/idem_report.md with four sections, ## 무엇이 잘못됐나, ## 어떻게 막았나, ## 남은 위험, and ## 운영 규칙 (the Korean headings mean: What went wrong, How it was blocked, Remaining risks, and Operating rules).
If you pick out requests, it does not prove anything. Send all of req_log in seq order and count the results as requests, created, replayed, rejected, transfers, total_transferred, v1_total, and double_paid_avoided. If you import idem_api.py as a module, you do not have to start a process 120 times. In the report, write as a number the amount you blocked this time.