The order arrived twice, and once it vanished
The one that vanished — commit timing, at-most-once and at-least-once
One-line summary
The point at which a consumer saves its position determines the semantics. If it saves before processing, a message whose processing died midway never comes again (at-most-once, "vanished"); if it saves after processing, a message that died before the save comes again (at-least-once, "arrived twice"). The answer is to choose the latter and make the processor idempotent, and rewinding (--reset-offsets) is the tool for recovering what vanished.
Why this was needed
The 'message delivery semantics' section of the design document covers the consumer side in two paragraphs. If a consumer reads a message, saves its position, and then processes it, then when it dies after saving but before processing, the process that takes over starts from the saved position, so the messages before it are never processed: at-most-once. If it reads, processes, and then saves, then when it dies after processing but before saving, the process that takes over receives again a message that was already processed: at-least-once. And the document adds: in many cases the message has a primary key, so the update is idempotent (receiving the same message twice only overwrites the same record).
These two paragraphs are the two incidents in the course title. The "one that vanished" is the result of saving before processing, and fixing it turns it into "arrived twice." The design is not to pick one of the two but to choose arriving twice and make it harmless to process twice.
How it works
Auto-commit runs regardless of processing. In the consumer configuration document, enable.auto.commit defaults to true, and the offset is committed in the background every auto.commit.interval.ms (default 5000). A processor that relies on auto-commit, like the console consumer, commits what was "read," not what was "processed." If a consumer handed over three records and committed offset 3, but the processor died on the second, the next consumer reads from 3. The second and third are still in Kafka, but this group passes over them forever.
shipments P0: order-1 order-2 order-3 order-4 order-5
ship-svc 읽음 3건 → 커밋 3 → 처리기 order-2 에서 크래시
처리됨: order-1 사라짐: order-2, order-3
What vanished is not the messages but the position. Consuming is not deleting, so if you read from the beginning with another group, all five are there. And with --reset-offsets in the operations document, you can rewind the group's position. There are scenarios such as --to-earliest, --to-latest, --to-offset, --shift-by, and --to-datetime; it only actually changes things if you add --execute (without it, it just shows the plan), and the consumer instances must be stopped. When you rewind, the two records that vanished come back, but the already processed order-1 also comes again. The price of rewinding is duplicates.
The idempotent consumer. This builds, on the processor's side, the case the document called "there is a primary key, so the update is idempotent." Record the processed order number in the same place as the result, and skip an incoming order if it is already there. If you keep the state only in memory, it dies with the process and the next run produces duplicates again. Whether it is a file or a DB, it has to survive together with the processing result.
Where does a new group start? The auto.offset.reset entry is about what to do when the group has no committed offset or that offset no longer exists (because the data was deleted). earliest goes to the very front, latest (the default) goes to the very back, by_duration:<ISO8601> goes back by that period from now, and none throws an exception. Because the default is latest, if you create a new group and attach it with no settings, it skips all the earlier messages, which is the cause of the report "the new service never saw the old orders." The console tool's --from-beginning is a shortcut that turns this into earliest, and it is ignored if the group already has an offset.
Commits disappear. In the broker configuration document, offsets.retention.minutes (default 10080, 7 days): if the group stays empty or stops subscribing to the topic and this time passes, the committed offsets are discarded. When a consumer comes back after that, it has no offset, so auto.offset.reset applies. If a batch consumer that stopped for a week comes back and starts with latest, it skips all the data in between.
What it looks like in practice
When you get a report that "orders vanished," the order is this. Describe the group to see the committed offsets, read that range with another group (or without a group) to confirm whether the messages exist, and if they do, it is a position problem: with the consumer stopped, rewind with --reset-offsets --to-offset and reprocess. Whether reprocessing creates duplicates depends on whether the processor is idempotent, and if it is not, you must fix that before rewinding.
Conversely, a report that "the same order was processed twice" is usually the normal behavior of at-least-once after a rebalance or restart. If you treat it as a bug and move the commit earlier, the next report becomes "it vanished." The two reports are the two ends of the same knob.
What you will do in the next lab
You reproduce the shipping processor dying on the second order so that two events vanish, confirm with another group that they are still in Kafka, rewind the group and reprocess to see duplicates appear, and block them with an idempotent processor. Finally, you compare the two values of auto.offset.reset using new groups.