The order arrived twice, and once it vanished
Retries create a second arrival — stop it with the idempotent producer
This lab runs on a VM
Apache Kafka 4.3.1 runs as a single KRaft node on an Ubuntu VM (at localhost:9092). The helper kafka-lab-slow on|off delays only requests of 4 KB or more to the broker by 600 ms, triggering producer timeouts and retries. It takes about 4 minutes to come up the first time.
Goal
You reproduce the same record being left in the log several times when a producer retries because it did not get a response, and confirm with kafka-dump-log.sh that an idempotent producer reduces it to one copy. You also look at the place idempotence does not cover (retransmission across producer sessions).
Why it matters
As the design document says, when a producer hits a network error, it cannot tell whether that error happened before or after the message was committed. So it sends again, and if the original request had succeeded, it is left in the log twice: that is at-least-once. An idempotent producer prevents this by having the broker give each producer an ID, attaching a sequence number to each record, and filtering out the same sequence number. In 4.x it is on by default (the default of enable.idempotence is true), but it is silently turned off if you touch one setting wrong, and it does not cover retransmission across sessions in the first place. If you have seen these three with your own eyes, you will know where to look when you get a report of "arrived twice."
Steps
- Create the topic
paymentswith 1 partition and put in two lines (pay-1,pay-2) with default settings. - Using
kafka-dump-log.sh, dump the segment ofpayments-0and save it to/root/kafka/dump-default.txt. The batch'sproducerIdmust not be -1 andbaseSequencemust be 0 (the default producer is idempotent). - With
enable.idempotence=false, put in one more line (pay-3), dump again, and save it to/root/kafka/dump-noidem.txt. The new batch must haveproducerId: -1. - Run
kafka-lab-slow on, create the topicpayments-dup, and put in a single 5,000-character line (/root/kafka/big.txt) with idempotence off, usingrequest.timeout.ms=300,delivery.timeout.ms=2000,retries=2, andmax.block.ms=10000. Keep the producer's stderr in/root/kafka/dup-producer.log, and when it finishes, runkafka-lab-slow off. The log must contain two or more copies of the same record. - Under the same conditions, put into the topic
payments-idemwith idempotence on (enable.idempotence=true) and keep stderr in/root/kafka/idem-producer.log. Retries happen, but only one copy must remain in the log. - With slow turned off, in the topic
payments-app, putpay-77by running the producer twice separately (with idempotence left on). Save the dump to/root/kafka/two-sessions.txt; theproducerIdof the two batches must differ. - In
/root/kafka/producer-report.txt, write five lines:dup_copies=<payments-dup 의 레코드 수>,idem_copies=<payments-idem 의 레코드 수>,two_sessions_copies=<payments-app 의 레코드 수>,acks_default=all, andidempotence_default=true. (In the first three, put the number of records in the named topic.)
Notes
- Dump:
kafka-dump-log.sh --files /var/lib/kafka/<토픽>-0/00000000000000000000.log --print-data-log(with the topic name in the placeholder).producerIdandbaseSequenceappear on the batch line, andpayloadon the record line. - Record count: the last number (the log-end offset) of
kafka-get-offsets.sh --bootstrap-server localhost:9092 --topic <토픽>. - Pass client settings with
--command-property 키=값(a key=value pair; in 4.3,--producer-propertyis deprecated). Make the 5,000-character line withhead -c 5000 /dev/zero | tr '\0' x > big.txt; echo >> big.txt. - Producer configuration document:
delivery.timeout.msmust be at leastrequest.timeout.ms + linger.ms, and to turn on idempotence you needacks=all,retries>0, andmax.in.flight.requests.per.connection<=5. If you explicitly turn idempotence on and give conflicting values, you get a ConfigException. - Common mistake 1: finishing step 4 with slow on and not turning it off. Small requests are unaffected so it is hard to notice, but the next steps that handle large records slow down.
- Common mistake 2: putting two lines through one producer in step 6. That is the same session, so the sequence numbers continue and it becomes a single batch. It has to be "run twice separately" to get two sessions.
Two records with the default producer
Create the topic payments with 1 partition and put in two lines (pay-1, pay-2) with default settings.
After --create --topic payments --partitions 1, run printf 'pay-1\npay-2\n' | kafka-console-producer.sh .... You can see the record count with kafka-get-offsets.sh.
The default producer is idempotent
Using kafka-dump-log.sh, dump the segment of payments-0 and save it to /root/kafka/dump-default.txt. The batch's producerId must not be -1 and baseSequence must be 0 (the default producer is idempotent).
kafka-dump-log.sh --files /var/lib/kafka/payments-0/00000000000000000000.log --print-data-log. The ID the broker gave the producer and the record sequence number are written as they are in the batch header; this is the material used to filter out duplicates.
With idempotence off, there is no ID
With enable.idempotence=false, put in one more line (pay-3), dump again, and save it to /root/kafka/dump-noidem.txt. The new batch must have producerId: -1.
--command-property enable.idempotence=false. A producer with idempotence off does not receive an ID, so the broker has no way to recognize it if the same record arrives again.
A retry creates the second arrival
Run kafka-lab-slow on, create the topic payments-dup, and put in a single 5,000-character line (/root/kafka/big.txt) with idempotence off, using request.timeout.ms=300, delivery.timeout.ms=2000, retries=2, and max.block.ms=10000. Keep the producer's stderr in /root/kafka/dup-producer.log, and when it finishes, run kafka-lab-slow off. The log must contain two or more copies of the same record.
If the producer gets no answer within 300 ms, it sends the same batch again (retries), and the broker writes both the original request that arrived late and the retry that arrived late. The producer ends up reporting failure, yet the log has three copies. Count them with kafka-get-offsets.sh.
An idempotent producer leaves only one copy
Under the same conditions, put into the topic payments-idem with idempotence on (enable.idempotence=true) and keep stderr in /root/kafka/idem-producer.log. Retries happen, but only one copy must remain in the log.
Turn on slow and send exactly as in step 4, changing only enable.idempotence=true. The retried batch arrives carrying the same sequence number, so the broker filters it out. REQUEST_TIMED_OUT is still printed in the producer's log: the retries did happen, and only the duplicates are absent.
If the session differs, idempotence does not cover it
With slow turned off, in the topic payments-app, put pay-77 by running the producer twice separately (with idempotence left on). Save the dump to /root/kafka/two-sessions.txt; the producerId of the two batches must differ.
Run printf 'pay-77\n' | kafka-console-producer.sh ... twice. With different processes, the broker gives a new producer ID and the sequence also starts from 0, so the broker sees different records. This is the shape of an application sending again after a failure, and what stops it is not an idempotent producer but a business key such as the order number.
Summarize with the record counts of the three topics
In /root/kafka/producer-report.txt, write five lines: dup_copies=<payments-dup 의 레코드 수>, idem_copies=<payments-idem 의 레코드 수>, two_sessions_copies=<payments-app 의 레코드 수>, acks_default=all, and idempotence_default=true. (In the first three, put the number of records in the named topic.)
The three numbers are the last field of kafka-get-offsets.sh. The other two are the defaults from the producer configuration document. The grader recounts the three numbers on the broker right now and compares them.