Production Backend API Capstone
A Retry Is Not a New Command
One-line summary
Even if a client resends the same request because it never received a response, the server must recognize it as the same command. Combining a transaction with an organization-scoped idempotency key means an order is created only once, and a retry returns the resource that was created first.
Why a timeout creates a duplicate order
If the network drops right after the server commits, the client does not know it succeeded. Retrying unconditionally creates a second order. The advice "send the request only once" cannot be honored in a distributed system. Instead, the client sends an Idempotency-Key and the server stores it together with the organization. The first write and the key record must be in the same transaction, so there is no state in which only one of them succeeds.
PostgreSQL's INSERT ... ON CONFLICT (org_id, idempotency_key) merges racing requests at the database's serialization point. Code that reads first and inserts later has a race condition: two requests both see "not found" at the same time and both insert. The conflict path must not overwrite the existing order with the newly submitted amount. It must return the original order so that the result of the same command stays stable.
How to handle failures in practice
Close driver connections and cursors with a context manager, and put writes inside an explicit transaction boundary. Guarantee a rollback when an exception occurs, and do not put SQL arguments or connection strings in the exception message. An integration test sends the same key twice with different amounts and confirms that the first response is 201, the retry is 200, and the ID and original amount are identical. A different organization must be able to use the same key independently.
Three mechanisms that make retries safe
Retries are unavoidable in a distributed system. The problem is that the caller cannot tell whether it is safe to retry. Not receiving a response is different from the work not having been done.
The requester creates the idempotency key. If the server created it, it would differ on every retry and be useless. The client generates one UUID and sends the same value from the first attempt to the last retry.
create table payment_requests(
org_id bigint not null,
idem_key uuid not null,
request_sha bytea not null,
status text not null check(status in ('처리중','완료','실패')),
response jsonb,
created_at timestamptz not null default now(),
primary key (org_id, idem_key));
Reject the same key with different content. That is why request_sha is stored alongside it. If the key is the same but the amount differs, that is a bug rather than a retry, and quietly letting it succeed is the worst outcome.
Claim the row first, then do the work. If insert ... on conflict do nothing returning id returns nothing, someone has already started. In that case, look at the state of that row: if it is complete, return the stored response as is; if it is in progress, return 409 so the caller asks again shortly. Making the server wait piles up those connections and creates another problem.
Do not bundle outside calls and the database transaction into one unit. If you keep a transaction open while calling a payment gateway, locks are held and a connection is tied up for as long as that slow call takes. Split the order: record the intent in a short transaction, call the outside service outside the transaction, and write the result in another short transaction. If the process dies in the middle, a row with only the intent is left behind, so build a recovery job alongside it that finds such rows and asks the other side for their status.
Always give retries a cap and jitter. If everyone retries at the same interval, the system collapses again the moment the outage recovers. Mix randomness into exponential backoff, set a total attempt time, and after that treat the request as failed so a person sees it. A queue that retries forever is a device for hiding outages.
Practical judgment criteria
Idempotency is a stronger contract than "duplicates produce an error." The user must be able to safely get the result again. In a real service you also state explicitly the key retention period, the policy for request-body hash collisions, and cleanup of old keys. This capstone first proves the most important creation boundary and conflict semantics, and the next module exposes them externally as actual HTTP statuses and body contracts.