TT Lab
Get started
Learn Learning paths Courses

The order arrived twice, and once it vanished

What to keep and what to drop — retention, compact, min.insync.replicas

Continue in TT Lab

One-line summary

Topic settings answer three questions. How long do we keep data (retention.ms, retention.bytes)? Do we keep only the last value for each key (cleanup.policy=compact)? How many replicas must confirm before a write counts as successful (min.insync.replicas and acks=all)? Deletion and compaction happen per segment and do not touch the active segment; most cases of "I deleted it but it's still there" are this.

Why this was needed

Kafka does not delete data when it is consumed, so something has to delete it. In the topic configuration document, the cleanup.policy entry provides two policies. delete (the default) discards old segments that have reached the retention time or size limit, and compact turns on log compaction, which keeps the latest value for each key. If you list both (delete,compact), old segments are discarded by the retention rules and the remaining segments are compacted. An empty list means infinite retention.

The reason compaction is needed is explained by way of example in the 'Log Compaction' section of the design document. When user 123's email has changed three times, time-based retention throws away old changes wholesale, so that even reading from the beginning you cannot reconstruct the current state. Compaction always keeps the last update for each key, so the log becomes a snapshot of the final value of every key; database change subscriptions, event sourcing, and state journaling stand on top of this.

How it works

Time-based retention is per segment. The default of retention.ms is 604800000 (7 days), and -1 means unlimited. The document calls this "an SLA on how quickly consumers must read." retention.bytes is a per-partition size limit and defaults to -1 (none). But the unit that gets deleted is not a record but a segment file. An old segment becomes a candidate for deletion only after the segment rolls upon reaching segment.ms (default 7 days) or segment.bytes (default 1 GiB), and the active segment is never deleted. On top of that, the broker, according to the broker configuration, checks only every log.retention.check.interval.ms (default 300000, 5 minutes). So even if you set retention.ms=1분 (that is, a retention of one minute), the data remains if the segment has not rolled or the check interval has not come. The lab VM has the check interval shortened to 10 seconds.

events (retention.ms=20000, segment.ms=10000)
  세그먼트 0: e1 e2 e3        ← 12초 뒤 e4 가 들어오며 굴러간다 (닫힘)
  세그먼트 4: e4              ← 활성. 절대 지워지지 않는다
  20초 + 검사 주기 뒤 → 세그먼트 0 삭제 → 가장 앞 오프셋이 3 이 된다

Compaction keeps the last value for each key. The guarantees in the design document: compaction does not change the order and only removes records, and offsets never change (a removed offset is treated as being at the same position as the next offset). A consumer that reads from the beginning sees the final state of every key in the order written. A record that writes a null value for a key is a deletion marker (tombstone); it causes the earlier records of that key to be removed, and it itself disappears after delete.retention.ms (default 1 day). When compaction happens is decided by min.cleanable.dirty.ratio (default 0.5: it moves only when half of the log is duplicates), min.compaction.lag.ms, and max.compaction.lag.ms, and here too the active segment is not a target. The lab lowers the dirty ratio to 0.01 so that compaction happens soon.

Records that are too large. max.message.bytes (default 1048588) is the maximum size of a record batch the broker accepts. A batch over it is rejected to the producer with a RecordTooLargeException, and nothing is left in the log; we confirmed this by measurement.

min.insync.replicas and acks=all. The min.insync.replicas entry is the minimum number of ISR members (including the leader) that must confirm for a write to succeed when a producer sends with acks=all. If the ISR is smaller than this number, the producer gets a NotEnoughReplicas or NotEnoughReplicasAfterAppend exception. The typical configuration the document gives is a replication factor of 3, min.insync.replicas=2, and acks=all: a majority must confirm the write for it to succeed, and regardless of acks, it is replicated to the entire ISR and is not visible to consumers until this condition is met. The default is 1, so there is no guarantee. The VM in this course has a single broker, so it cannot reproduce this interaction (by measurement too, no error occurred on a single node), and we check it only through the quiz. unclean.leader.election.enable (default false) is whether to elect a replica outside the ISR as leader as a last resort, and turning it on can lose data.

Changing at runtime. 'Modifying topics' in the operations document: run kafka-configs.sh --entity-type topics --entity-name X --alter --add-config k=v to add a setting (in the Korean source this command is broken across lines, which leaves the stray span 로 설정을 더하고, meaning "to add a setting"), and use --delete-config k to remove one. There is no need to recreate the topic.

What it looks like in practice

The report "I reduced retention to 1 day but the disk didn't shrink" is resolved by looking at the segment size. A partition where segment.bytes is 1 GiB and 200 MB accumulates per day only rolls a segment after five days, and until then nothing is deleted. The answer is to shrink segment.ms as well.

For "it's a compacted topic but I see old values," check three things. Is it in the active segment (it is not compacted)? Has it passed the dirty ratio (default 0.5)? And is the consumer reading from the beginning and has not yet reached the head (compaction happens at the tail)? All three are normal behavior.

What you will do in the next lab

On a topic with 20-second retention, you roll the segment and watch the earliest offset rise; on a compacted topic you watch alice shrink to a single last value; you confirm that a large record is rejected by max.message.bytes; and then you change the retention of a running topic.