Order Status Sometimes Goes Backwards — Build the Receiving Side
Goal
You build the receiving side of webhooks that another system pushes in. You attach HMAC signature verification with constant-time comparison, a time window that blocks replays, deduplication by the two keys of delivery id and event id, absorption of order reversal by version comparison, and a structure that gives a 200 quickly and processes later, then replay a day's deliveries and tally each branch.
Why it matters
A webhook is not something we call but something we receive, so control is on the other side. The receiving endpoint must be open, so anyone can POST to it, if our response is late the other side resends, and order is not guaranteed. So the receiving side must judge four things for itself. Who sent it (signature), when it was sent (time window), whether it was already seen (duplicate), and whether it is newer than the current one (version). If even one is missing, fake events get into the ledger or the order state goes backward. The fact that the duplicate key is two is especially a trap. The delivery id refers to one transmission, and the event id refers to one occurred event. What must be blocked is duplicate application of an event, so filtering by the delivery id alone blocks only half. The grader does not believe your sentences. It starts the sender and receiver you built directly on ports the grader chooses, and reruns the verifier and ledger with secrets and deliveries the grader created to check the answers.
Steps
- Create /root/wh/sender.py, run it on port 8012, and save
/deliveriesto /root/wh/deliveries.json and the shared secret to /root/wh/secret.txt. - Create /root/wh/verify.py so that it compares the signature in constant time and outputs
{"ok": ..., "reason": ...}. - Attach a time window to verify.py so that it rejects deliveries outside the window as
stale. - Create /root/wh/ledger.py so that it filters duplicates with the two keys of delivery id and event id.
- Make ledger.py apply only when it is newer than the currently stored version, and record a late-arriving old version as
stale_version. - Create /root/wh/receiver.py so that it only verifies, puts the item in a queue, and then gives a 200 right away, and so that
--drainflushes the queue into the ledger. - With /root/wh/replay_day.py, replay all of the day's deliveries and create /root/wh/day.db and /root/wh/result.json.
- Report in four sections in /root/wh/wh_report.md.
Notes
- Sender run contract:
python3 /root/wh/sender.py --port <포트> [--secret <비밀>](the placeholders are the port and the secret)./healthreturns{"ok": true, "events": 30, "deliveries": 41}, and/deliveriesreturns{"now": <기준 시각>, "tolerance": 300, "deliveries": [...]}(the placeholder is the reference time). One delivery is{"delivery_id": ..., "signature": ..., "body": <원본 문자열>}(the placeholder is the original string). - Signature format:
t=<epoch>,v1=<hex>. The signature material is"<t>.<body>"and it is HMAC-SHA256. Use the body as the received string. If you parse and serialize it again, the signature goes out of sync. - This day's data is 41 deliveries: 30 events (10 orders × 3 versions) plus 4 resends, 3 new deliveries of the same event, 2 old replays, and 2 fake signatures. The order in which the versions arrive differs per order.
- Verifier run contract:
python3 verify.py --secret <파일> --delivery <파일> [--now <epoch>] [--tolerance <초>](the placeholders are files and seconds) outputs{"ok": true|false, "reason": "ok"|"bad_signature"|"stale"|"malformed"}. If you do not pass--now, it uses the current time. Accept all four arguments in step 2 too (you attach the window in step 3). If it is not a signature shape it is malformed, if the signature is wrong it is bad_signature, and if the signature is right but the time is outside the window it is stale. - Ledger run contract:
python3 ledger.py --db <sqlite> --delivery <파일>(the placeholders are the sqlite file and the delivery file) outputs{"stored": ..., "applied": ..., "reason": ...}. reason is new, duplicate_delivery, duplicate_event, or stale_version. The tables must includeorder_state(order_id, version, status). - Receiving endpoint run contract:
python3 receiver.py --port <포트> --db <sqlite> --secret <파일> [--tolerance <초>](the placeholders are port, sqlite, file, seconds) providesGET /healthandPOST /webhook. The delivery id comes in theX-Delivery-Idheader and the signature in theX-Signatureheader. If it passes, 200{"queued": true}, and if caught at the signature or time window, 400. If you pass--drain, it does not start the server but flushes the queue into the ledger and then outputs the tally. The queue table name isinbox. - Replayer run contract:
python3 replay_day.py --deliveries <파일> --db <sqlite> --out <파일>(the placeholders are files). The tally slots are ten: deliveries, accepted, rejected_signature, rejected_stale, duplicate_delivery, duplicate_event, stored, applied, stale_version, and orders.acceptedis the deliveries that passed the signature and time window, andstoredis the number of events that entered the ledger because they were not duplicates. - Common mistakes: parsing the body and serializing it again to compute the signature, comparing signatures with
==, filtering duplicates by the delivery id only, and processing all the way to the ledger at the receiving point so the 200 is late. - Run the server in the background, wait until
/healthis 200, and then move on. The grader does not look at the process you left running but restarts the scripts directly.
Get a day's deliveries in hand
Create /root/wh/sender.py, run it on port 8012, and save the /deliveries response to /root/wh/deliveries.json and the shared secret to /root/wh/secret.txt. There are 41 deliveries and 30 events.
Make the two routes /health and /deliveries with flask. Build the delivery list by adding resends, new deliveries of the same event, old replays, and fake signatures to the 30 events. If you include the reference time in the response, the test will not be shaken by the clock later.
Tell who sent it by the signature
Create /root/wh/verify.py so that it verifies the signature of one delivery and outputs {"ok": ..., "reason": ...}. If it is not a signature shape, it is malformed, and if the signature is wrong, it is bad_signature. You must compare in constant time.
Split the signature string t=...,v1=... to get t and v1, and compute HMAC-SHA256 with "<t>.<body>" as the material. The body must be used as the received string. Python's hmac module has a function that compares in time proportional to the length — == ends at the first differing byte, so the time leaks the secret.
The hand that pushes an old one back in
Attach a time window to verify.py. Even if the signature is right, if the difference between --now and the delivery's t exceeds --tolerance (default 300 seconds), it must be rejected as stale. Both the past and the future side are outside the window.
A signature alone cannot stop pushing a previously exchanged valid request back in as is. So the signature material contains a timestamp, and the receiving side looks at the difference from the current time. You must block the future side too — the other side's clock being fast and someone writing t into the future cannot be distinguished. The signature comes first in the order.
There are two keys
Create /root/wh/ledger.py so that it puts deliveries that passed verification into the ledger, but filters as duplicate_delivery if the same delivery id comes again, and as duplicate_event if the delivery id is new but the event id already exists. The tables must be kept in a sqlite file.
The approach of checking and inserting if absent lets two arriving in the same instant both pass. INSERT directly into tables that have each of the two keys as the primary key and use the constraint violation exception as the signal for "already exists." The delivery id check comes first — if you swap the order, a resend gets recorded as duplicate_event.
So the order state does not go backward
Make ledger.py keep order_state(order_id, version, status) and apply only when it is newer than the currently stored version. An old version that arrives late has stored true but applied false, and the reason is stale_version.
You must not trust the arrival order. Compare the version the event carries with the currently stored value. If it is an order seen for the first time, put it in as is, and if it already exists, update only when it is larger. A case where the same version comes again is not an update target either.
Give a 200 quickly and process later
Create /root/wh/receiver.py so that POST /webhook looks only at the signature and time window, puts the item in the inbox queue, and then gives a 200 {"queued": true} right away. It must not touch the ledger at the receiving point. --drain does not start the server but flushes the queue into the ledger.
The delivery id comes in the X-Delivery-Id header and the signature in the X-Signature header. Do not parse the body; pass the original string as is to verification. If caught in verification, it is 400. If you separate putting in the queue from applying to the ledger, the processing time no longer touches the other side's timeout.
The whole day again, without dropping any
With /root/wh/replay_day.py, replay all the deliveries of /root/wh/deliveries.json and create /root/wh/day.db and /root/wh/result.json. The tally slots are ten: deliveries, accepted, rejected_signature, rejected_stale, duplicate_delivery, duplicate_event, stored, applied, stale_version, and orders.
Use the now and tolerance contained in the delivery list file as the reference time as they are. That way the test is not shaken by the clock. Those that failed verification do not go on to the ledger, and those caught as duplicates are not counted in stored. Replay all of them without dropping any.
Receiving-side inspection report
In /root/wh/wh_report.md, write four sections, ## 무엇이 들어왔나, ## 중복을 어떻게 걸렀나, ## 순서를 어떻게 다뤘나, and ## 남은 위험과 운영 규칙 (in order: what came in, how we filtered duplicates, how we handled order, remaining risks and operating rules). The numbers from result.json must be in the body.
The reader is both our team lead and someone at the partner company. Write the count for each branch as a number, and also write what would have been missed if you had filtered by the delivery id alone. The remaining risks must include secret rotation and the width of the time window.