The order arrived twice, and once it vanished
The second arrival — duplicates from retries and the idempotent producer
One-line summary
When a producer does not get a response, it has no way to know whether the message was committed, so it sends again, and if the original request actually succeeded, two copies remain in the log. An idempotent producer filters out the retransmission using the producer ID and record sequence number given by the broker, leaving only one copy. In 4.x it is on by default, but it does not cover retransmission across sessions.
Why this was needed
The 'message delivery semantics' section of the design document states the problem precisely. If a producer hits a network error while publishing, it cannot tell whether that error happened before or after the message was committed; it is like a connection dropping during an INSERT into a table with an auto-generated key. Before 0.11, a producer had no choice but to send again, which made it at-least-once. If the original request actually succeeded, the retransmission writes the same message to the log one more time.
Starting with 0.11, an idempotent delivery option exists. The broker gives each producer an ID, the producer attaches a sequence number to every message it sends, and the broker filters out messages with the same ID and sequence number. Transactions were introduced around the same time, making it possible to write atomically to multiple partitions. This course covers only idempotence; transactions are a layer on top of it.
How it works
acks decides what to wait for. The acks entry in the producer configuration document: 0 does not wait for any confirmation from the server at all, so there is no guarantee it was received and no retries happen (the offset is always -1). 1 answers once the leader writes to its own log, so if the leader dies before the followers replicate, the message is lost. all (= -1) waits until the entire ISR has confirmed, and it is the strongest guarantee: nothing is lost as long as at least one member of the ISR stays alive. The default is all, and to turn on idempotence it must be all.
Retries are practically infinite by default. The default of retries is 2147483647, and the document advises you to leave that value alone and govern the total time of retries with delivery.timeout.ms (default 120000). delivery.timeout.ms must be at least request.timeout.ms (default 30000) plus linger.ms. The request.timeout.ms entry has this sentence: this value must be larger than the broker's replica.lag.time.max.ms (default 30000) to reduce the chance of message duplication caused by unnecessary retries. This is the place where the document names duplication explicitly as a result of retries.
Idempotence has conditions. The enable.idempotence entry: when it is on, each message is written to the stream exactly one time, and when it is off, retries caused by broker failures and the like can write duplicates. To turn it on, max.in.flight.requests.per.connection must be 5 or less, retries must be greater than 0, and acks must be all. If there is a conflicting setting and idempotence was not explicitly turned on, idempotence is silently turned off. If you turn it on explicitly and there is a conflict, you get a ConfigException. So the moment you put in acks=1 "for performance," duplicate protection disappears without any warning.
멱등 프로듀서의 배치 헤더 (kafka-dump-log.sh)
producerId: 1 producerEpoch: 0 baseSequence: 0 lastSequence: 1 ← ID 와 순번
멱등을 끈 배치
producerId: -1 producerEpoch: -1 baseSequence: -1 lastSequence: -1 ← 걸러 낼 재료가 없다
The broker recognizes a retransmission from this header. In a measurement (4.3.1, with a 600 ms delay injected into a local broker, request.timeout.ms=300, retries=2), the producer with idempotence off left three copies of the same record, and the producer with idempotence on left only one; both producers ended up reporting "failure." They simply did not receive the response; the record had been on the broker.
What idempotence does not cover. The ID and sequence number belong to the producer session. When a process starts anew, it gets a new ID and the sequence starts from 0, so when an application is told of a failure and calls send again, the broker sees a new record. The transactional.id entry calls this "reliability across multiple producer sessions" and places it in the domain of transactions. Without that layer, the answer is to filter on the consumer side by a business key (the order number).
What it looks like in practice
When a report comes in on a payment service that "the same payment was recorded twice," the first thing to check is the producer configuration. If there is a value like acks=1 or max.in.flight=10, idempotence is silently off. The second check is the batch header in the log: if the same payment is in two batches with different producerId values, it is an application-layer retransmission, and this cannot be prevented by producer configuration.
Conversely, "it says it was sent but it isn't there" is the shape of acks=0. The producer reports success the moment it puts the record in the socket buffer, and does not know whatever happens after that. That is an acceptable trade for latency-sensitive log collection, but not for payments.
What you will do in the next lab
You compare the batch headers of a default producer and an idempotence-off producer using a dump, turn on a slow network to reproduce retries leaving three copies and idempotence reducing them to one, and then run the producer twice separately to see that retransmission across sessions is not blocked.