TT Lab
Get started
Learn Learning paths Courses

System Integration (EAI)

"Exactly Once" Is a Lie

Continue in TT Lab

Summary

"Exactly once" is not something the delivery layer gives you but is built from at-least-once delivery + deduplication on the receiver, and the latter half is our responsibility.

Why use a queue

Synchronous integration holds only if the other side is alive. Asynchronous removes that premise.

In exchange there is a price. Complexity and duplicates. "We cannot get the response immediately" means you have to design a result notification channel separately, and "we retry" means the same message may be processed twice.

Three grades of delivery guarantee

Grade Meaning Reality
at-most-once At most once. Loss possible Things that can be lost, like logs and metrics
at-least-once At least once. Duplicates possible Most practical messaging
exactly-once Exactly once Holds only conditionally

What you must understand here.

"Exactly once" = "at-least-once delivery" + "deduplication on the receiver"

In other words, exactly-once is not something the delivery layer makes by magic but something the receiver makes appear so by filtering out duplicates. Even if a broker advertises "exactly-once support," that is within specific conditions (the same cluster, using the transaction API). It breaks the moment you write to an external system.

Therefore the consumer must always be written assuming duplicates. It is the default, not the exception.

The scope of ordering guarantees

Another thing often misunderstood.

Order is guaranteed only within a partition (or queue). Global ordering across the whole topic is not guaranteed.

So if there is a requirement that "messages with the same order number must be processed in order," you must use the order number as the partition key. Then messages of the same order go into the same partition and the order is kept.

If you do not know this and distribute round-robin, a cancel message can be processed before the order message. And that happens only once or twice a day, so it is very hard to find the cause.

The standard shape of a consumer design

1. 메시지 수신
2. 파싱 및 형식 검증        → 실패: 즉시 error/DLQ (재시도해도 똑같다)
3. 멱등 확인 (이미 처리했나) → 이미 처리: 아무것도 안 하고 ack
4. 업무 처리 (DB 트랜잭션)
5. 처리 이력 기록           ← 4와 같은 트랜잭션 안에서
6. ack (큐에서 제거)

The key is putting step 5 in the same transaction as step 4. If done separately, you get a state of "the business was processed but the history could not be recorded," and retrying in that state results in duplicate processing.

And you must not do step 6 before step 4. If you ack first and die while processing, the message disappears. This is the point where it becomes at-most-once.

Do not retry parse failures

There is a reason step 2 is separate. A message with broken JSON or a missing required field fails 100 times if retried 100 times. But if it gets caught in the retry logic, that message keeps failing at the head of the queue and blocks the normal messages behind it. This is called a poison message.

So errors are divided into two kinds.

Without this distinction the queue gets blocked, and a blocked queue is a total business stop.

Queue lag is the most important metric

If you want to see the health of asynchronous integration as a single number, it is the consumer lag. "The number of accumulated messages" or "the age of the oldest unprocessed message."

Set alerts not on absolute values but on trend and duration. "Lag increasing for 10 consecutive minutes" is a better condition than "lag over 1000." In a system with batch-like inflow, a large momentary lag is normal.

A file queue is a queue too

Many SI sites have no broker (Kafka, RabbitMQ). In that case you use a directory-based queue. It is surprisingly robust.

/data/if/inbox/       ← 도착
/data/if/processing/  ← 처리 중 (원자적 mv 로 이동 = 잠금)
/data/if/done/        ← 성공
/data/if/error/       ← 실패 (DLQ 역할)

The key is that mv is atomic within the same filesystem. Only the process that succeeds in the inbox → processing move gets that message. Even if you start several consumers, there is no duplicate processing.

Two points to note.

  1. mv between different filesystems is copy + delete and not atomic. You must move within the same mount.
  2. A consumer can pick up a file before the sender has finished writing it. So you write under a temporary name and rename after completion, or use a convention of also creating a completion flag file (.ok).

Build the reprocessing procedure in advance

When an outage occurs you will certainly have to reprocess. What you need then.

You must write this procedure as a document before launch and rehearse it. If you make it on the day of the outage, you stay up all night.

What it looks like in the field

What actually blows up in a project that introduced a queue is not the queue itself but the consumer-side assumptions.

The most common is duplicates. When the network drops once and reconnects, the broker resends messages it has not yet received an acknowledgment for. This is not a fault but behavior per the specification, and if the consumer did not know that, the same order is loaded twice. And this incident is usually discovered because the amount does not match in month-end settlement — three weeks after the incident.

The second is ordering. If you build assuming the whole queue preserves order, it breaks the moment you add partitions or consumers. The range where ordering is guaranteed is usually inside a single partition, so you must choose the key so that messages of the same order go to the same partition.

The third is lag. Because a queue absorbs load, the sender notices nothing wrong even if the consumer is slow. So if you do not monitor lag, you learn that it has been backed up for hours from an inquiry like "I cannot see today's data."