FDE Capstone: The Warehouse Got the Same Order Three Times
The response vanished, but the order survived
In one line
A request failure and a business failure are not the same. You retry with the same business key, and you can tell a duplicate reservation apart only by reconciling the response with the ledger.
Why this was needed
You sent a reservation to the warehouse API. A connection error appeared on the screen. Is it safe to send again? If the connection dropped before the reservation was saved, you have to send again for the reservation to exist. Conversely, if the saving was finished and only the response disappeared, the moment you send again with no protection, two reservations can be created. The error sentences seen from the client are similar, but the business state inside the server is the opposite.
This fictional warehouse actually produces this difference. The first request of ORD-BEFORE returns 503 before saving. The first request of ORD-AFTER commits the reservation to SQLite and then closes the connection with no HTTP response. If you handle the two situations with just one failure flag, you design the recovery wrongly. That is why, instead of ending with reading the logs, you run the program you built directly through both situations.
How it works
The submission API is POST /reservations. Only order_id, sku, and quantity go in the JSON. A separate Idempotency-Key header identifies the same business attempt. In this course the order number is used as the key. The API and the client jointly promise the following three cases.
| Request | The warehouse's judgment | Response |
|---|---|---|
| A key seen for the first time and a valid body | Save a new reservation | 201, reservation ID |
| The same key and the same body | Return the existing reservation | 200, existing reservation ID |
| The same key and a changed body | Conflicts with the existing business | 409, no new reservation |
If you create a new key every time, every retry looks like a new business. Even if you store the processed keys only in an in-memory set, the memory disappears when the process restarts. This server stores the key, the body, and the reservation in the same SQLite transaction and uses a uniqueness constraint on the key. With an application's "look up first and write if absent" alone, two requests that come in at the same time can both see an empty state. The database has to hold the final duplicate judgment as well.
What if you use a hash of the whole body as the key? The moment the quantity of the same order changes, the hash changes too. The server sees this as a different key and creates a new reservation. That is not conflict resolution but sidestepping the conflict check. If the customer wants to allow order amendments, an amendment API and a version rule have to be agreed separately. In this exercise, you hold the 409 and keep the existing reservation.
The client has a 1-second limit per request and allows at most four POST attempts, only for 503 and connection errors. It waits 0.1 seconds between retries. A 409 or a contract error stops, because waiting does not change the meaning of the request. These values are a policy for a short teaching verification. In production, you decide them separately, taking into account the other side's limits, the latency distribution, the overall processing deadline, and the load. It is not a recommendation of four attempts for every service.
You do not record confirmed just because you received a response. You look up the business result with GET /orders/, and also check that it is exactly one reservation, that the SKU and quantity are right, and, if there is an ID received from the POST, that it is the same as the looked-up ID. If there are three reservations, you must not pick only the first and write success. If the lookup fails, even if you received an ID once, you have not met this contract's confirmation condition, so you leave unconfirmed and null.
The status names are also distinguished. confirmed means the result was confirmed by the set criteria, and unconfirmed means it has not been confirmed yet. rejected means the request was rejected with a 409 or a contract error. If you translate unconfirmed as "no reservation", the next owner risks resending with a new key. What you could not observe must be left as not observed.
Why was it reserved only once when it was sent twice
Let us follow ORD-AFTER in time order. On the first POST, the server stores the key and the body and creates a reservation ID. Then it closes the connection, so the ID does not reach the client. The second POST has the same key and the same body, so it does not create a new reservation and returns the existing ID. The last GET checks that there is only one reservation for that order. The report says attempts=2, but the actual ledger has one reservation. The two numbers differing is not an error; it is because retries and business effects were counted separately.
ORD-BEFORE also does the POST twice, but at the first request the ledger is empty. The reservation is created for the first time at the second request. If you look only at the final number of reservations, the two cases are the same, so an error-injection test also has to distinguish when the saving happened. On the other hand, a wrong answer that changes the retry key may happen to work well for ORD-BEFORE but creates two reservations for ORD-AFTER. This is why you must not say you verified the recovery strategy from the success of a single case.
What it looks like in the field
Reprocessing the same file is not an unexpected accident but a common business flow. When the person submitting changes or the job is interrupted, the same file comes in again. So the verification also does not look only at the first run. You restart the server process but keep the same ledger, send the same input again, and check that not only the number of reservations but also the IDs are kept. Checking the business invariant answers this question more directly than "it exited without errors".
This exercise is a single SQLite ledger of a single warehouse. It does not solve distributed transactions, inventory consistency across several warehouses, or redelivery after a key expires. Leaving these limits in the handoff document is also part of the implementation.
What to check next
In the next quiz, you judge a lost response, a restart, and a content change as different situations. When you implement too, if you keep the error kind, the actual number of attempts, and the business result as separate variables, it is easier to explain what you know and what you do not know.