Not a Call We Make, but One We Receive
Summary
A webhook is an integration where control is on the other side, so the receiving side must judge for itself four things: who sent it (signature), when it was sent (time window), whether it was already seen (duplicate), and whether it is newer (order), and it must return a 200 before it has finished all those judgments.
Why this was needed
In an integration where we call an API, we decide the start time and the retry policy. A webhook is the opposite. The other side POSTs to our address, and we can neither refuse nor postpone that request. Three things bite us at the same time here.
First, anyone can POST to our address. A webhook receiving endpoint must be open to the internet for the other side to call. That means others can call it too. Nobody knowing the address is not a defense.
Second, the other side sends again. If we give the 200 late or cannot give it, the other side sees a failure and resends. That is, the longer our processing time, the more duplicates increase. And a resend does not come only when we failed to process. Much more often we had already processed it and only the response was late, so the other side did not receive it.
Third, order is not guaranteed. Just because three events, accepted → paid → shipped, occurred for one order does not mean they arrive in that order. If the other side sends in parallel, or the next one arrives first while one is going through retries, the order is reversed. If you overwrite in arrival order, the order state goes backward.
How it works
Signature. The widely used method is to send an HMAC computed with a shared secret in a header. RFC 2104 defines HMAC, and practical implementations generally follow the shape shown in Stripe's webhook documentation and GitHub's delivery validation documentation. The header carries a timestamp t and a signature v1, and the material of the signature is a string concatenating the timestamp and the original body.
The most common mistake here is to compute the signature after parsing the body and serializing it again. If the key order or whitespace differs by even one character, the signature becomes an entirely different value. Signature verification must be done on the bytes exactly as received.
The second mistake is the comparison method. If you compare with got == want, it ends right at the first differing byte, and that time difference gives an attacker the information "the first few characters were right." In Python, hmac.compare_digest compares in time proportional to the length.
Time window. Even if the signature is right, it does not mean it was sent now. Pushing a previously exchanged valid request back in as is is called a replay, and a signature alone does not stop it. So you put a timestamp in the signature material, and the receiving side checks whether the difference from the current time is within a window (usually a few minutes). If you narrow the window, normal requests are rejected with even a small clock drift, and if you widen it, the replayable interval gets longer.
Duplicates. The trap is that there are two keys. The delivery id refers to one transmission, and the event id refers to one occurred event. If it comes again with the same delivery id, it is a resend, and if the delivery id is new but the event id is the same, that is also the same event. What must be blocked is duplicate application of an event, so the real key is the event id. If you filter by the delivery id alone, you miss the latter.
Order. Do not trust the arrival order; judge by the version or occurrence time the event carries. Apply only when it is newer than what is currently stored, and quietly drop the old one.
POST /webhook
│
├─ 서명 틀림 ──────────▶ 400 (원장에 손대지 않는다)
├─ 시각 창 밖 ─────────▶ 400
└─ 통과 ─▶ 큐에 적재 ─▶ 200 (여기서 끝. 처리는 뒤에서)
│
└─ 배수(drain) ─▶ 중복인가? 더 새것인가? ─▶ 원장 적용
Giving the 200 quickly is the last piece. If you touch even the ledger at the receiving point, the slower the processing, the more it hits the other side's timeout, then resends increase, and as resends increase, processing gets even slower. If you only verify, put it in a queue, and answer right away, this loop is broken.
What it looks like in the field
First, "sometimes the order status goes backward" is the most common report. The cause is almost always overwriting in arrival order. Looking at the log, after shipped, paid has been applied.
Second, signature verification quietly breaks in frameworks. In a framework that parses the body automatically, you must find a separate way to get the original bytes. There are also cases where a proxy rewrites the body (decompression, charset conversion).
Second and a half, secret rotation is not in the plan. The moment you change the secret, every delivery sent before it is rejected. So during the rotation period accept both the old and the new secret and pass the request if either one matches.
Third, you do not decide what to drop when the queue backs up. Webhooks keep coming in, so the queue grows endlessly. If each event has a version, old events for the same target may be dropped, but that judgment must be decided in advance.
Fourth, do not trust it even if the documentation says the other side preserves order. One retry on their side breaks the order. If the receiving side has a version comparison, there is no loss, and if not, an incident occurs.
What you will do in the next lab
You start a partner sender that pushes order events and get a day's 41 deliveries in hand. Among them are mixed resends, new deliveries of the same event, old replay attempts, and fakes made by someone who does not know the secret. You build a signature verifier that compares in constant time, attach a time window, filter duplicates with the two keys of delivery id and event id, and apply only when the version is newer. Then you build a receiving endpoint that gives a 200 quickly and puts the item in a queue, and finally you replay the whole day without dropping any and tally how many there were in each branch.