The order arrived twice, and once it vanished
Same key, same partition — touch the log by hand
This lab runs on a VM
Apache Kafka 4.3.1 runs as a single KRaft node on an Ubuntu VM (the systemd kafka service, at localhost:9092). /opt/kafka/bin is on the PATH, so you can use tools such as kafka-topics.sh right away. It takes about 4 minutes to come up the first time.
Goal
You create a topic and put keyed events into it, and see that the same key goes to the same partition and that order is preserved within a partition. You read a consumer group's offsets and lag, confirm that two groups read independently of each other, and then observe for yourself that increasing the number of partitions changes the key placement.
Why it matters
To read an incident like "an order arrived twice and one vanished," you first need to know what Kafka promises and what it does not. Kafka promises only ordering within a partition. Because the same key goes to the same partition, the order of events for one order is preserved, but across different orders it is not. And a consumed message is not deleted: because each consumer group has its own integer (the offset) for "how far have I read," what one group missed another group can read again from the beginning. These two facts are the basis of the duplication and loss stories in the later modules.
Steps
- Create the topic
orderswith 3 partitions. - Put the nine lines of
/root/kafka/orders.txt(in키:값format, meaning key:value) into it withkafka-console-producer.sh, parsing the key. They must end up spread over several partitions. - Read everything from the beginning so that partition, offset, and key are visible, and save it to
/root/kafka/keyed.txt. The same key must appear in only one partition. - Read nine events from the beginning with the consumer group
order-svc, and save the result of describing that group to/root/kafka/group.txt. - Add the three lines of
/root/kafka/more.txt(without reading them), describeorder-svcagain, and save it to/root/kafka/lag.txt. LAG must be visible. - Read twelve events from the beginning with a second group,
analytics. The offsets oforder-svcmust stay the same. - Increase
ordersto 6 partitions, addorder-1:refunded, read from the beginning, and save it to/root/kafka/repartition.txt. Then write four lines to/root/kafka/repartition-report.txt:partitions_before=3,partitions_after=6,order1_before=<3단계에서 order-1 이 있던 파티션>, andorder1_after=<refunded 가 들어간 파티션>. (For order1_before, put the partition where order-1 was in step 3; for order1_after, put the partition that received refunded.)
Notes
- Key parsing:
--reader-property parse.key=true --reader-property key.separator=:(in 4.3,--propertyis deprecated). - Showing metadata when reading:
--formatter-property print.key=true --formatter-property print.partition=true --formatter-property print.offset=true. The output has the formPartition:0\tOffset:0\t키\t값(the last two columns are the key and the value). - A consumer finishes after reading N events with
--from-beginning --max-messages N --timeout-ms 8000. Without them you have to press Ctrl-C. - Group status:
kafka-consumer-groups.sh --bootstrap-server localhost:9092 --describe --group <이름>(put the group name in the placeholder). After the consumer finishes, "has no active members" appears, but the offset table is still shown. - Increasing partitions:
kafka-topics.sh ... --alter --topic orders --partitions 6. Decreasing is not supported (operations document). - Common mistake 1: without
--max-messages, giving only--timeout-msmakes it wait that long after the last message. Give both. - Common mistake 2: expecting old data to be moved after changing the number of partitions. Kafka does not redistribute existing data.
A topic with 3 partitions
Create the topic orders with 3 partitions.
kafka-topics.sh --bootstrap-server localhost:9092 --create --topic ... --partitions 3. After creating it, check PartitionCount with --describe.
Put in keyed events
Put the nine lines of /root/kafka/orders.txt (in 키:값 format, meaning key:value) into it with kafka-console-producer.sh, parsing the key. They must end up spread over several partitions.
Pass --reader-property parse.key=true --reader-property key.separator=: and feed the file in as standard input. If you do not parse the key, the whole line becomes the value and the key is null, so the partition is chosen regardless of the key.
See partitions and offsets with your own eyes
Read everything from the beginning so that partition, offset, and key are visible, and save it to /root/kafka/keyed.txt. The same key must appear in only one partition.
To --from-beginning --max-messages 9 --timeout-ms 8000, add --formatter-property print.key=true --formatter-property print.partition=true --formatter-property print.offset=true and send the output to a file. The order is mixed across partitions, but within one partition the offsets go up.
A consumer group's offsets
Read nine events from the beginning with the consumer group order-svc, and save the result of describing that group to /root/kafka/group.txt.
If you pass --group order-svc, the position you read to is committed to the broker. When it finishes, CURRENT-OFFSET in kafka-consumer-groups.sh --describe --group order-svc should equal LOG-END-OFFSET for each partition and LAG should be 0.
If you do not read, lag accumulates
Add the three lines of /root/kafka/more.txt (without reading them), describe order-svc again, and save it to /root/kafka/lag.txt. LAG must be visible.
Put them in the same way as in step 2, but do not run a consumer. LOG-END-OFFSET goes up and CURRENT-OFFSET stays the same, so the difference is the LAG.
Groups are independent of each other
Read twelve events from the beginning with a second group, analytics. The offsets of order-svc must stay the same.
--group analytics --from-beginning --max-messages 12. Consumed messages are not deleted, so a new group can read everything from the beginning, and it has no effect at all on the offsets of other groups.
Increasing partitions changes the key placement
Increase orders to 6 partitions, add order-1:refunded, read from the beginning, and save it to /root/kafka/repartition.txt. Then write four lines to /root/kafka/repartition-report.txt: partitions_before=3, partitions_after=6, order1_before=<3단계에서 order-1 이 있던 파티션>, and order1_after=<refunded 가 들어간 파티션>. (For order1_before, put the partition where order-1 was in step 3; for order1_after, put the partition that received refunded.)
After --alter --topic orders --partitions 6, put in one line with key parsing and read with --max-messages 13 as in step 3. As the operations document says, the default partitioner is hash(key) % number of partitions, so when the number of partitions changes the same key can go to a different partition, and existing data is not moved.