The Myth of Exactly Once
Summary
Exactly-once = at-least-once delivery + deduplication on the receiving side. Do not wait for the infrastructure to do it by magic; make the consumer idempotent.
Why this was needed
You pressed the pay button and the screen froze. The user presses it again. The server received two requests, and if it processes both, that is a double charge.
The root of this scene is the indistinguishability we saw earlier. When the sender has not received an acknowledgement, there is no way to know whether the message failed to arrive or arrived and only the response was lost. So to prevent loss you must retry, and retrying creates duplicates.
There are only two strategies. At-most-once — send and forget, loss possible, no duplicates. At-least-once — retry until an acknowledgement is received, no loss, duplicates possible. Most systems choose the latter, because loss is worse than duplication.
How it works
First, you need to distinguish idempotent operations from non-idempotent ones. "Set the email to a@b.com" gives one result no matter how many times you send it. "Increase the balance by 100" gives 200 if you send it twice. Setting is idempotent and incrementing is not.
HTTP already specifies this at the method level. GET/HEAD/OPTIONS are safe and idempotent, PUT/DELETE are idempotent, and POST is not. PUT is idempotent because it is a set operation, "make this resource have this value". DELETE is the same — even if you are told to delete something already deleted, the final state is the same. The response code may differ, to 404, but idempotency is a property of state, not of response codes.
The standard way to make a POST idempotent is an idempotency key. The client generates one UUID per request and puts it in a header, and the server remembers that key and, when the same key arrives again, does not process it anew but returns the stored first response as is. Stripe and PayPal use exactly this approach.
There are four things to watch in the implementation.
First, storing and expiring keys. You cannot keep them forever, so you usually set 24 hours.
Second, concurrency. If two requests with the same key arrive almost simultaneously, both may mistake it for a "key seen for the first time" and process twice. You must claim the key atomically at the moment you receive it. With Redis, a single SET key <state> NX EX 86400 is enough.
Third, verifying that the key and the body match. If the key is the same but the body differs, it is a client mistake or an attack. It is safer to reject it.
Fourth, distinguishing the nature of the failure. For a transient failure (DB timeout), the retry must really be attempted again, but for a definitive failure (insufficient balance), it is better to cache it and give the same answer.
What you meet in the field
In message queue consumers, the key often already exists. It is an event ID or an aggregate ID + sequence number. You just record the processed ID in a Redis set or a DB unique constraint. If you use a unique constraint, the database also guarantees atomicity for you, so the race condition disappears.
One trap to watch for. If "recording that it was processed" and "actually processing it" are in different stores, it is again a dual-write problem. If possible, put them in the same transaction.
Where do failed messages go
Even with idempotency, a problem remains. Messages that fail forever. The format is broken, the data they reference was deleted, or there is a defect in the code, so no matter how many times you try again the same exception occurs. If such a message is at the head of the queue, all the normal messages behind it are blocked, and retries keep consuming resources.
So you need two mechanisms.
Put a cap and an interval on retries. If you retry immediately, it fails again for the same reason, so you gradually lengthen the interval. And if several consumers failed at the same time, their retry times are also clustered, so you mix in a little randomness to scatter them. Without this, the moment the other service recovers it gets hit by a retry storm and collapses again.
Send it to a dead-letter queue once it exceeds the cap. It is a place that takes the message out of the processing flow without discarding it. When you move it here, leave along with the original message why it failed, how many times it was tried, and what the last error was. Without this information, even if you open that queue later, you cannot tell what to do.
A dead-letter queue is easy to create and forget. If you do not set an alert on the length of that queue, nobody looks at it. And you also need to decide what to do when the alert fires. Whether to put the messages back after fixing the code, fix the data, or whether they can simply be discarded. This judgment usually belongs to the business owner, so there must also be a way to show what failed in a form a person can read.
Finally, one point about processing order. Retries break ordering. This is because while an earlier message is deferred by a retry, a later message is processed first. If ordering is important in the flow, you must either process messages with the same key in a single line, or design each message to carry the final state so that you do not depend on order at all. The latter is much sturdier, and it is the same story as the earlier "setting is idempotent and incrementing is not".
What you will do in the next lab
You build a payment service that deliberately double-charges, attach an idempotency key, add atomic claiming so it is processed only once even under concurrent requests, reject body mismatches, and even set key expiration.