TT Lab
Get started
Learn Learning paths Courses

Idempotency — Two Clicks, One Charge

Retries Are Unavoidable

Continue in TT Lab

In one line

Retries cannot be avoided. So the server has to prevent duplicates.

Why it is needed: the client cannot tell the difference

You sent a payment request, but no response came. Which is it?

  1. The request did not arrive at the server → you must send it again
  2. It arrived and was processed, but only the response did not come back → if you send it again, the payment is made twice

The client has no way to tell these apart. So whichever one it picks, it is wrong.

So the client has to retry, and the only option is to make the server recognize the second one.

Idempotency key

For each request, the client creates and attaches a unique key.

POST /pay
Idempotency-Key: 9f2c-...-a1
{"user":"u1","amount":1000}

When retrying, it uses the same key. If the server has seen that key before, it does not process the request and returns the stored response as it is.

This is the heart of it: the second request also gets a success response. It is not an error. From the client's point of view, it is "requested once, succeeded once," and that is correct.

Where to store it

An in-memory dictionary will not do. There are two reasons.

  1. It forgets on restart
  2. There are several servers: what Pod 1 remembers, Pod 2 does not know

The second is far more common. Test with one Pod and it works perfectly, and the moment autoscaling kicks in, duplicates appear.

So the storage has to be a place everyone looks at together: a DB or Redis.

When the same key comes with a different body

Idempotency-Key: k1   {"amount": 1000}
Idempotency-Key: k1   {"amount": 99999}

If you silently let the second one through, you return a 1000-won response to a 99999-won request. The client thinks 99999 won was charged.

So you store the hash of the request body together with the key, and if it differs, reject it with 422. It is a device to keep a bug that reuses keys from being silently buried.

When the same key arrives at the same time

This is the place that goes wrong most often.

row = db.get(key)          # 없다
if not row:                # ← 열 개가 동시에 여기를 통과한다
    charge()
    db.put(key, response)

Someone else cuts in between the lookup and the insert. The naive approach breaks the moment load arrives.

To do it properly, put a unique constraint on the key and have the side whose insert fails wait and read the stored response, or wrap it in a lock. Using the guarantee the database already has is the cheapest way.

How long to keep the stored response

You cannot keep it forever. Usually you keep it for about 24 hours and then delete it.

If it is too short, a late retry creates a duplicate, and if it is too long, the storage keeps growing. Set it comfortably longer than the client's retry window.

Which errors to retry

Response Retry Why
No response (timeout) ✅ You do not know whether it arrived
500, 502, 503 ✅ A temporary problem on the server side
429 ✅ (after waiting) Honor Retry-After
400, 422 ❌ The request is wrong. It stays the same forever
404 ❌ Usually permanent

And do it with growing intervals (exponential backoff), and shake it per client (jitter). Otherwise the retries pile onto a recovering server at the same moment and knock it down again.

Which requests to attach it to

Attach it to requests that move money or create something. POST, and PATCH when it has side effects.

GET·PUT·DELETE are methods designed to be idempotent from the start. With PUT, writing the same thing twice gives the same result, and with DELETE, the second one is "already gone." That is the design; the implementation does not become that way by itself.

Traps when returning the stored response

Once you add idempotency, new problems arise. Returning the original response as it is is not always right.

The state may have changed in the meantime. Suppose a user cancels an order after creating it, and then a retry arrives and you return the stored 201 response ("Created") as it is: the client believes a live order exists. What you store must be the result of that request, and if the client needs to know the current state, give a lookup path along with the response.

If you store the whole response, it grows. If you keep large bodies as they are, the storage fills quickly. What you need is usually only the status code and the identifier of the created resource, so store only that and recreate the rest.

Decide a retention period and delete. An idempotency key is not needed forever. Set it comfortably longer than the client's retry window (usually 24 hours) and delete after that. If you do not delete, the table keeps growing, and as the index grows, lookups that were originally fast get slow.

delete from idempotency_keys where created_at < now() - interval '7 days';

Decide what to do when an expired key comes back. If you silently process it as new, a duplicate is created, and if you reject it, a very late retry fails. Rejecting is safer, and in that case state clearly "this key has expired, so send it again with a new key."

Tell whether it is a retry with a response header. If you attach a header such as Idempotency-Replayed: true, both the client and whoever investigates know whether this response was newly processed or stored. If you also leave it in the log, "requests doubled" is immediately separated into retries or a real increase.

Do not trust the key itself. It is a value the client made, so it may collide with another user's key. Uniqueness must be established together with an organization or user identifier, so that nobody ends up receiving someone else's response.

In the field

A duplicate payment report is usually discovered first not by the user but in settlement. The user remembers only "it seemed the payment did not go through so I pressed once more," and what happened in between is left only in the server logs.

That is why an idempotency key is unusually hard to add after an incident. Clients that have already gone out do not send keys, so for a while the server has to accept requests with keys and without keys together. If you do not decide beforehand what to do with requests that have no key (let them through as they are, reject them, or allow a grace period for a set time), a bigger mess arises in the middle of the migration.

If there are clients whose releases you cannot force, such as mobile apps, there is also a method where the server builds a temporary key from the hash of the request body and prevents duplicates only for a short time. It is not complete, but the most common case, where a user presses twice in a row, gets filtered out.