TT Lab
Get started
Learn Learning paths Courses

The order arrived twice, and once it vanished

Retention and compaction — what topic configs keep and drop

Continue in TT Lab

This lab runs on a VM

Apache Kafka 4.3.1 runs as a single KRaft node on an Ubuntu VM (at localhost:9092). The broker's retention check interval (log.retention.check.interval.ms) is shortened to 10 seconds and the cleaner backoff (log.cleaner.backoff.ms) to 5 seconds, so you can see deletion and compaction within a minute. It takes about 4 minutes to come up the first time.

Goal

You see for yourself that time-based retention (retention.ms) deletes old segments, that compaction (cleanup.policy=compact) keeps only the last value for each key, and that max.message.bytes rejects a large record, and then you change the settings of a running topic.

Why it matters

Topic settings decide "what is kept and for how long." The default retention is 7 days, and because deletion is per segment, data does not disappear right away just because retention.ms has passed: the active segment is never deleted, and it can be deleted only after the segment rolls (segment.ms). Compaction is a different kind of retention. In a keyed change log it always keeps the last value of each key, so reading from the beginning restores the current state. If you do not know this difference, you cannot read the reports "I deleted it but it's still there" and "it was supposed to be kept but it vanished."

Steps

  1. Create the topic events with retention.ms=20000 and segment.ms=10000.
  2. Put in e1, e2, and e3, wait at least 12 seconds, then put in e4 to roll the segment, and wait about another 30 seconds. Once the earliest offset is no longer 0, write two lines to /root/kafka/retention.txt: earliest=<가장 앞 오프셋> and latest=<로그 끝 오프셋> (the earliest offset and the log-end offset).
  3. Create the topic customers with cleanup.policy=compact, segment.ms=10000, and min.cleanable.dirty.ratio=0.01.
  4. With key parsing, put in alice:v1, bob:v1, alice:v2, and alice:v3, wait at least 12 seconds, then put in carol:v1 to roll the segment. When compaction has finished (about 30 seconds), read from the beginning with the keys and save it to /root/kafka/compacted.txt. alice must appear only once, with v3.
  5. Create the topic photos with max.message.bytes=1024 and try putting in a single 2,000-character line (/root/kafka/big2k.txt). Keep the producer's stderr in /root/kafka/toolarge.txt; it must be rejected, leaving 0 records.
  6. With kafka-configs.sh, for events, change retention.ms to 86400000 (without recreating the topic).
  7. In /root/kafka/topic-report.txt, write five lines: events_retention_ms=86400000, customers_policy=compact, photos_max_bytes=1024, min_insync_default=1, and three_broker_min_isr=2.

Notes

A topic that keeps only 20 seconds

Create the topic events with retention.ms=20000 and segment.ms=10000.

--create --topic events --partitions 1 --config retention.ms=20000 --config segment.ms=10000. Without segment.ms, the segment does not roll for the default 7 days and nothing is deleted.

The old segment is deleted

Put in e1, e2, and e3, wait at least 12 seconds, then put in e4 to roll the segment, and wait about another 30 seconds. Once the earliest offset is no longer 0, write two lines to /root/kafka/retention.txt: earliest=<가장 앞 오프셋> and latest=<로그 끝 오프셋> (the earliest offset and the log-end offset).

If kafka-get-offsets.sh --topic events --time earliest changes to something like events:0:3, the first segment has been deleted. The unit that is deleted is not a record but a segment, so e1–e3 disappear all at once and e4 in the active segment remains.

A topic that keeps only the last value per key

Create the topic customers with cleanup.policy=compact, segment.ms=10000, and min.cleanable.dirty.ratio=0.01.

--config cleanup.policy=compact --config segment.ms=10000 --config min.cleanable.dirty.ratio=0.01. The default dirty ratio is 0.5, so the cleaner moves only when half of the log is duplicates; in the lab we lower it to almost 0.

After compaction there is one alice

With key parsing, put in alice:v1, bob:v1, alice:v2, and alice:v3, wait at least 12 seconds, then put in carol:v1 to roll the segment. When compaction has finished (about 30 seconds), read from the beginning with the keys and save it to /root/kafka/compacted.txt. alice must appear only once, with v3.

The active segment is not compacted, so you have to roll it with carol. If you pass --formatter-property print.key=true when reading, the form is alice\tv3. If there are still three alice records, the cleaner has not run yet, so read again after 10 seconds.

A record that is too large is rejected

Create the topic photos with max.message.bytes=1024 and try putting in a single 2,000-character line (/root/kafka/big2k.txt). Keep the producer's stderr in /root/kafka/toolarge.txt; it must be rejected, leaving 0 records.

head -c 2000 /dev/zero | tr '\0' y > big2k.txt; echo >> big2k.txt. When the broker rejects it because of the batch size limit, a RecordTooLarge-type error is printed in the producer's log and kafka-get-offsets.sh returns 0.

Change settings while running

With kafka-configs.sh, for events, change retention.ms to 86400000 (without recreating the topic).

kafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name events --alter --add-config retention.ms=86400000. It is exactly 'Modifying topics' in the operations document. Verify with --describe.

Summarize with the setting values

In /root/kafka/topic-report.txt, write five lines: events_retention_ms=86400000, customers_policy=compact, photos_max_bytes=1024, min_insync_default=1, and three_broker_min_isr=2.

For the first three, read the current values with kafka-configs.sh --describe and write them (the grader looks at it the same way). The last two come from the min.insync.replicas entry in the topic configuration document: the default, and the typical value when the replication factor is 3.