The order arrived twice, and once it vanished
Retention and compaction — what topic configs keep and drop
This lab runs on a VM
Apache Kafka 4.3.1 runs as a single KRaft node on an Ubuntu VM (at localhost:9092). The broker's retention check interval (log.retention.check.interval.ms) is shortened to 10 seconds and the cleaner backoff (log.cleaner.backoff.ms) to 5 seconds, so you can see deletion and compaction within a minute. It takes about 4 minutes to come up the first time.
Goal
You see for yourself that time-based retention (retention.ms) deletes old segments, that compaction (cleanup.policy=compact) keeps only the last value for each key, and that max.message.bytes rejects a large record, and then you change the settings of a running topic.
Why it matters
Topic settings decide "what is kept and for how long." The default retention is 7 days, and because deletion is per segment, data does not disappear right away just because retention.ms has passed: the active segment is never deleted, and it can be deleted only after the segment rolls (segment.ms). Compaction is a different kind of retention. In a keyed change log it always keeps the last value of each key, so reading from the beginning restores the current state. If you do not know this difference, you cannot read the reports "I deleted it but it's still there" and "it was supposed to be kept but it vanished."
Steps
- Create the topic
eventswithretention.ms=20000andsegment.ms=10000. - Put in
e1,e2, ande3, wait at least 12 seconds, then put ine4to roll the segment, and wait about another 30 seconds. Once the earliest offset is no longer 0, write two lines to/root/kafka/retention.txt:earliest=<가장 앞 오프셋>andlatest=<로그 끝 오프셋>(the earliest offset and the log-end offset). - Create the topic
customerswithcleanup.policy=compact,segment.ms=10000, andmin.cleanable.dirty.ratio=0.01. - With key parsing, put in
alice:v1,bob:v1,alice:v2, andalice:v3, wait at least 12 seconds, then put incarol:v1to roll the segment. When compaction has finished (about 30 seconds), read from the beginning with the keys and save it to/root/kafka/compacted.txt.alicemust appear only once, withv3. - Create the topic
photoswithmax.message.bytes=1024and try putting in a single 2,000-character line (/root/kafka/big2k.txt). Keep the producer's stderr in/root/kafka/toolarge.txt; it must be rejected, leaving 0 records. - With
kafka-configs.sh, forevents, changeretention.msto86400000(without recreating the topic). - In
/root/kafka/topic-report.txt, write five lines:events_retention_ms=86400000,customers_policy=compact,photos_max_bytes=1024,min_insync_default=1, andthree_broker_min_isr=2.
Notes
- Topic settings are given at creation with
--config 키=값(a key=value pair), and later withkafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name <토픽> --alter --add-config 키=값(the topic name goes in the placeholder). To verify, use--describeof the same command. - Offset range:
kafka-get-offsets.sh --topic events --time earliestand--time latest. The output has the form토픽:파티션:오프셋(topic:partition:offset). - Checking compaction:
--from-beginning --timeout-ms 8000 --formatter-property print.key=true. Compaction changes neither the order nor the offsets (a guarantee in the design document);alice v3stays in its original place (offset 3). - The last two lines are conceptual. The default of
min.insync.replicasin the topic configuration document is 1, and the typical configuration the document gives is replication factor 3 with min.insync.replicas 2 and acks=all. On a single node you cannot reproduce that interaction, so we check it with the quiz. - Common mistake 1: waiting without rolling the segment. The active segment is neither deleted nor compacted;
segment.msis checked only when a new record comes in. - Common mistake 2: treating it as a failure when
alicestill appears several times after compaction. The cleaner has not run yet, so read again after 10 seconds.
A topic that keeps only 20 seconds
Create the topic events with retention.ms=20000 and segment.ms=10000.
--create --topic events --partitions 1 --config retention.ms=20000 --config segment.ms=10000. Without segment.ms, the segment does not roll for the default 7 days and nothing is deleted.
The old segment is deleted
Put in e1, e2, and e3, wait at least 12 seconds, then put in e4 to roll the segment, and wait about another 30 seconds. Once the earliest offset is no longer 0, write two lines to /root/kafka/retention.txt: earliest=<가장 앞 오프셋> and latest=<로그 끝 오프셋> (the earliest offset and the log-end offset).
If kafka-get-offsets.sh --topic events --time earliest changes to something like events:0:3, the first segment has been deleted. The unit that is deleted is not a record but a segment, so e1–e3 disappear all at once and e4 in the active segment remains.
A topic that keeps only the last value per key
Create the topic customers with cleanup.policy=compact, segment.ms=10000, and min.cleanable.dirty.ratio=0.01.
--config cleanup.policy=compact --config segment.ms=10000 --config min.cleanable.dirty.ratio=0.01. The default dirty ratio is 0.5, so the cleaner moves only when half of the log is duplicates; in the lab we lower it to almost 0.
After compaction there is one alice
With key parsing, put in alice:v1, bob:v1, alice:v2, and alice:v3, wait at least 12 seconds, then put in carol:v1 to roll the segment. When compaction has finished (about 30 seconds), read from the beginning with the keys and save it to /root/kafka/compacted.txt. alice must appear only once, with v3.
The active segment is not compacted, so you have to roll it with carol. If you pass --formatter-property print.key=true when reading, the form is alice\tv3. If there are still three alice records, the cleaner has not run yet, so read again after 10 seconds.
A record that is too large is rejected
Create the topic photos with max.message.bytes=1024 and try putting in a single 2,000-character line (/root/kafka/big2k.txt). Keep the producer's stderr in /root/kafka/toolarge.txt; it must be rejected, leaving 0 records.
head -c 2000 /dev/zero | tr '\0' y > big2k.txt; echo >> big2k.txt. When the broker rejects it because of the batch size limit, a RecordTooLarge-type error is printed in the producer's log and kafka-get-offsets.sh returns 0.
Change settings while running
With kafka-configs.sh, for events, change retention.ms to 86400000 (without recreating the topic).
kafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name events --alter --add-config retention.ms=86400000. It is exactly 'Modifying topics' in the operations document. Verify with --describe.
Summarize with the setting values
In /root/kafka/topic-report.txt, write five lines: events_retention_ms=86400000, customers_policy=compact, photos_max_bytes=1024, min_insync_default=1, and three_broker_min_isr=2.
For the first three, read the current values with kafka-configs.sh --describe and write them (the grader looks at it the same way). The last two come from the min.insync.replicas entry in the topic configuration document: the default, and the typical value when the replication factor is 3.