TT Lab
Get started
Learn Learning paths Courses

The order arrived twice, and once it vanished

Same key, same partition — touch the log by hand

Continue in TT Lab

This lab runs on a VM

Apache Kafka 4.3.1 runs as a single KRaft node on an Ubuntu VM (the systemd kafka service, at localhost:9092). /opt/kafka/bin is on the PATH, so you can use tools such as kafka-topics.sh right away. It takes about 4 minutes to come up the first time.

Goal

You create a topic and put keyed events into it, and see that the same key goes to the same partition and that order is preserved within a partition. You read a consumer group's offsets and lag, confirm that two groups read independently of each other, and then observe for yourself that increasing the number of partitions changes the key placement.

Why it matters

To read an incident like "an order arrived twice and one vanished," you first need to know what Kafka promises and what it does not. Kafka promises only ordering within a partition. Because the same key goes to the same partition, the order of events for one order is preserved, but across different orders it is not. And a consumed message is not deleted: because each consumer group has its own integer (the offset) for "how far have I read," what one group missed another group can read again from the beginning. These two facts are the basis of the duplication and loss stories in the later modules.

Steps

  1. Create the topic orders with 3 partitions.
  2. Put the nine lines of /root/kafka/orders.txt (in 키:값 format, meaning key:value) into it with kafka-console-producer.sh, parsing the key. They must end up spread over several partitions.
  3. Read everything from the beginning so that partition, offset, and key are visible, and save it to /root/kafka/keyed.txt. The same key must appear in only one partition.
  4. Read nine events from the beginning with the consumer group order-svc, and save the result of describing that group to /root/kafka/group.txt.
  5. Add the three lines of /root/kafka/more.txt (without reading them), describe order-svc again, and save it to /root/kafka/lag.txt. LAG must be visible.
  6. Read twelve events from the beginning with a second group, analytics. The offsets of order-svc must stay the same.
  7. Increase orders to 6 partitions, add order-1:refunded, read from the beginning, and save it to /root/kafka/repartition.txt. Then write four lines to /root/kafka/repartition-report.txt: partitions_before=3, partitions_after=6, order1_before=<3단계에서 order-1 이 있던 파티션>, and order1_after=<refunded 가 들어간 파티션>. (For order1_before, put the partition where order-1 was in step 3; for order1_after, put the partition that received refunded.)

Notes

A topic with 3 partitions

Create the topic orders with 3 partitions.

kafka-topics.sh --bootstrap-server localhost:9092 --create --topic ... --partitions 3. After creating it, check PartitionCount with --describe.

Put in keyed events

Put the nine lines of /root/kafka/orders.txt (in 키:값 format, meaning key:value) into it with kafka-console-producer.sh, parsing the key. They must end up spread over several partitions.

Pass --reader-property parse.key=true --reader-property key.separator=: and feed the file in as standard input. If you do not parse the key, the whole line becomes the value and the key is null, so the partition is chosen regardless of the key.

See partitions and offsets with your own eyes

Read everything from the beginning so that partition, offset, and key are visible, and save it to /root/kafka/keyed.txt. The same key must appear in only one partition.

To --from-beginning --max-messages 9 --timeout-ms 8000, add --formatter-property print.key=true --formatter-property print.partition=true --formatter-property print.offset=true and send the output to a file. The order is mixed across partitions, but within one partition the offsets go up.

A consumer group's offsets

Read nine events from the beginning with the consumer group order-svc, and save the result of describing that group to /root/kafka/group.txt.

If you pass --group order-svc, the position you read to is committed to the broker. When it finishes, CURRENT-OFFSET in kafka-consumer-groups.sh --describe --group order-svc should equal LOG-END-OFFSET for each partition and LAG should be 0.

If you do not read, lag accumulates

Add the three lines of /root/kafka/more.txt (without reading them), describe order-svc again, and save it to /root/kafka/lag.txt. LAG must be visible.

Put them in the same way as in step 2, but do not run a consumer. LOG-END-OFFSET goes up and CURRENT-OFFSET stays the same, so the difference is the LAG.

Groups are independent of each other

Read twelve events from the beginning with a second group, analytics. The offsets of order-svc must stay the same.

--group analytics --from-beginning --max-messages 12. Consumed messages are not deleted, so a new group can read everything from the beginning, and it has no effect at all on the offsets of other groups.

Increasing partitions changes the key placement

Increase orders to 6 partitions, add order-1:refunded, read from the beginning, and save it to /root/kafka/repartition.txt. Then write four lines to /root/kafka/repartition-report.txt: partitions_before=3, partitions_after=6, order1_before=<3단계에서 order-1 이 있던 파티션>, and order1_after=<refunded 가 들어간 파티션>. (For order1_before, put the partition where order-1 was in step 3; for order1_after, put the partition that received refunded.)

After --alter --topic orders --partitions 6, put in one line with key parsing and read with --max-messages 13 as in step 3. As the operations document says, the default partitioner is hash(key) % number of partitions, so when the number of partitions changes the same key can go to a different partition, and existing data is not moved.