TT Lab
Get started
Learn Learning paths Courses

Lakehouse Table Format — Understanding Apache Iceberg Through Its Metadata

Undo means moving a pointer, not copying files — time travel, rollback, tags and branches

Continue in TT Lab

In one line

Because an Iceberg commit adds a new snapshot without deleting old ones, reading the past (time travel) means reading the list of an old snapshot, and undoing (rollback) means moving the main pointer to an old snapshot. A tag gives a point in time a name and a retention period, and a branch creates a lineage that grows separately from main.

Why it is needed — a DELETE at three in the morning

A DELETE that left out one condition wiped out all the orders of one region. In a traditional lake, you would have to find the files in a backup, copy them, and line them up by hand so they do not get mixed with the data that arrived in the meantime. It takes hours, and meanwhile the dashboards show wrong numbers.

In Iceberg, the DELETE that caused the accident is also a commit and merely created a new snapshot; the snapshot just before the accident and its files are still there. What you need to undo it is a single metadata commit that changes "which snapshot main points to".

How it works — reading the past and undoing

Spark's time travel comes in two forms.

SELECT count(*) FROM lake.demo.events VERSION AS OF 1234567890123456789;  -- 스냅샷 ID·브랜치·태그 이름
SELECT count(*) FROM lake.demo.events TIMESTAMP AS OF '2026-04-01 09:00:00';

It only reads the manifest list of that snapshot, so the cost is the same as reading the current table. There is one condition — that snapshot and its files must still exist.

For undoing, there are several procedures.

Procedure What it does
rollback_to_snapshot Rolls main back to an old snapshot in the current lineage
rollback_to_timestamp Rolls back to the snapshot that was current at that time
set_current_snapshot Sets a snapshot (or ref) that does not have to be an ancestor as current
cherrypick_snapshot Moves the changes of one snapshot onto the current state as a new snapshot (append and dynamic overwrite only)
fast_forward Advances one branch to the latest snapshot of another branch

A rollback does not create a new snapshot, and it does not delete the accident snapshot. One more line, "accident → just before the accident", is added to the path main has travelled (the snapshot-log). Removing the accident snapshot is the job of a later expiration.

Tags and branches — named references

The snapshot references of the spec keep refs in the metadata. Each ref has a snapshot ID and a type (tag or branch), and retention policies: max-ref-age-ms (the lifetime of the ref itself) and, for a branch, min-snapshots-to-keep and max-snapshot-age-ms. main is a branch that never expires.

A tag is a name label attached to one snapshot. You no longer need to memorize a 19-digit ID, and, more importantly, expiration does not delete the snapshot a tag points to. The Branching docs give an example of tagging weekly and monthly snapshots for auditing and giving them a retention period such as RETAIN 7 DAYS. According to the retention policy, expiration first removes refs whose max-ref-age-ms has passed, and keeps the snapshots that the remaining refs point to.

A branch is a lineage that grows separately from main. You commit new data to a branch rather than main first, and when the quality checks pass, you advance main to the head of that branch (fast_forward). Readers see only main, so they never see the data before the checks. The docs call this the audit branch (write-audit-publish, WAP), and show how to make existing jobs write to a branch without modifying them, using write.wap.enabled and spark.wap.branch. fast_forward works only when main is an ancestor of the branch, so it is rejected if another commit slipped into main in the meantime.

What it looks like in the field

Undoing after an accident. The on-call person looks at the snapshots main has travelled with SELECT * FROM 표.history (the placeholder stands for the table), reads the snapshot just before the accident with VERSION AS OF to check the numbers, and then calls rollback_to_snapshot. It takes a few seconds. The accident snapshot is kept so that the cause can be investigated.

Expiration ran first. If the cleanup job that runs every day deletes snapshots older than 5 days, by the time you realize you have to go back to the state of a week ago, it is already too late. For points in time you might have to go back to (month-end close, just before a deployment), put a tag in advance.

Unvalidated data went out to the dashboard. If the load job writes straight to main, it is already too late even when the quality check fails. If you write to a branch, check, and then publish, the failed data stays only on the branch.

What really matters in practice

What you will do in the next lab

You load three days, one commit per day, and then cause an accident on purpose by deleting one region. You read the row count just before the accident with VERSION AS OF, put a tag with 7-day retention on that snapshot, and roll main back with rollback_to_snapshot. Then you create a branch fix, write the next day's data only to the branch, and publish it by advancing main with fast_forward.