"Exactly Once" Is a Lie
Summary
"Exactly once" is not something the delivery layer gives you but is built from at-least-once delivery + deduplication on the receiver, and the latter half is our responsibility.
Why use a queue
Synchronous integration holds only if the other side is alive. Asynchronous removes that premise.
- Time decoupling — even if the other side is not there now, it is processed later
- Load absorption — even if 10,000 per second arrive, pile them in the queue and process 1,000 per second
- Failure isolation — the receiver's outage does not spread to the sender
- Reprocessing — failed messages can be put back in
In exchange there is a price. Complexity and duplicates. "We cannot get the response immediately" means you have to design a result notification channel separately, and "we retry" means the same message may be processed twice.
Three grades of delivery guarantee
| Grade | Meaning | Reality |
|---|---|---|
| at-most-once | At most once. Loss possible | Things that can be lost, like logs and metrics |
| at-least-once | At least once. Duplicates possible | Most practical messaging |
| exactly-once | Exactly once | Holds only conditionally |
What you must understand here.
"Exactly once" = "at-least-once delivery" + "deduplication on the receiver"
In other words, exactly-once is not something the delivery layer makes by magic but something the receiver makes appear so by filtering out duplicates. Even if a broker advertises "exactly-once support," that is within specific conditions (the same cluster, using the transaction API). It breaks the moment you write to an external system.
Therefore the consumer must always be written assuming duplicates. It is the default, not the exception.
The scope of ordering guarantees
Another thing often misunderstood.
Order is guaranteed only within a partition (or queue). Global ordering across the whole topic is not guaranteed.
So if there is a requirement that "messages with the same order number must be processed in order," you must use the order number as the partition key. Then messages of the same order go into the same partition and the order is kept.
If you do not know this and distribute round-robin, a cancel message can be processed before the order message. And that happens only once or twice a day, so it is very hard to find the cause.
The standard shape of a consumer design
1. 메시지 수신
2. 파싱 및 형식 검증 → 실패: 즉시 error/DLQ (재시도해도 똑같다)
3. 멱등 확인 (이미 처리했나) → 이미 처리: 아무것도 안 하고 ack
4. 업무 처리 (DB 트랜잭션)
5. 처리 이력 기록 ← 4와 같은 트랜잭션 안에서
6. ack (큐에서 제거)
The key is putting step 5 in the same transaction as step 4. If done separately, you get a state of "the business was processed but the history could not be recorded," and retrying in that state results in duplicate processing.
And you must not do step 6 before step 4. If you ack first and die while processing, the message disappears. This is the point where it becomes at-most-once.
Do not retry parse failures
There is a reason step 2 is separate. A message with broken JSON or a missing required field fails 100 times if retried 100 times. But if it gets caught in the retry logic, that message keeps failing at the head of the queue and blocks the normal messages behind it. This is called a poison message.
So errors are divided into two kinds.
- Permanent errors (format errors, missing required values, nonexistent codes) → DLQ immediately
- Transient errors (DB connection failure, the other system's 5xx, timeouts) → retry
Without this distinction the queue gets blocked, and a blocked queue is a total business stop.
Queue lag is the most important metric
If you want to see the health of asynchronous integration as a single number, it is the consumer lag. "The number of accumulated messages" or "the age of the oldest unprocessed message."
- Usually near 0 → normal
- Keeps increasing → the consumer cannot keep up (a performance problem or consumer down)
- Suddenly spikes → a surge on the producer side or a consumer outage
- A constant value that does not decrease → possibly blocked by a poison message
Set alerts not on absolute values but on trend and duration. "Lag increasing for 10 consecutive minutes" is a better condition than "lag over 1000." In a system with batch-like inflow, a large momentary lag is normal.
A file queue is a queue too
Many SI sites have no broker (Kafka, RabbitMQ). In that case you use a directory-based queue. It is surprisingly robust.
/data/if/inbox/ ← 도착
/data/if/processing/ ← 처리 중 (원자적 mv 로 이동 = 잠금)
/data/if/done/ ← 성공
/data/if/error/ ← 실패 (DLQ 역할)
The key is that mv is atomic within the same filesystem.
Only the process that succeeds in the inbox → processing move gets that message.
Even if you start several consumers, there is no duplicate processing.
Two points to note.
mvbetween different filesystems is copy + delete and not atomic. You must move within the same mount.- A consumer can pick up a file before the sender has finished writing it.
So you write under a temporary name and rename after completion,
or use a convention of also creating a completion flag file (
.ok).
Build the reprocessing procedure in advance
When an outage occurs you will certainly have to reprocess. What you need then.
- What failed — the messages in error/ and the reasons
- Why it failed — classification by reason (format/business/system)
- Can it be fixed — re-inject after fixing the data vs ask the source to resend
- Is reprocessing safe — is idempotency guaranteed
- Who approves — an interface where money moves may need approval
You must write this procedure as a document before launch and rehearse it. If you make it on the day of the outage, you stay up all night.
What it looks like in the field
What actually blows up in a project that introduced a queue is not the queue itself but the consumer-side assumptions.
The most common is duplicates. When the network drops once and reconnects, the broker resends messages it has not yet received an acknowledgment for. This is not a fault but behavior per the specification, and if the consumer did not know that, the same order is loaded twice. And this incident is usually discovered because the amount does not match in month-end settlement — three weeks after the incident.
The second is ordering. If you build assuming the whole queue preserves order, it breaks the moment you add partitions or consumers. The range where ordering is guaranteed is usually inside a single partition, so you must choose the key so that messages of the same order go to the same partition.
The third is lag. Because a queue absorbs load, the sender notices nothing wrong even if the consumer is slow. So if you do not monitor lag, you learn that it has been backed up for hours from an inquiry like "I cannot see today's data."