TT Lab
Get started
Learn Learning paths Courses

Lakehouse Table Format — Understanding Apache Iceberg Through Its Metadata

A catalog is an atomic pointer from one name to one metadata file

Continue in TT Lab

In one line

The catalog holds one slot, "table name → path of the current metadata.json", and it handles both the atomicity of commits and the refereeing of concurrent writes through a conditional swap that changes that slot only if it still holds the value that was read. Names, locations, and files are different layers.

Why a catalog was needed

As the previous module showed, the state of an Iceberg table is determined by a single metadata.json file. But a new metadata file is created with every commit. So someone has to tell you "which file is the current one?", and when two writers commit at the same time, someone has to accept only one of them.

Early on, the file system did this job. The file system tables approach in the spec tries a rename to a fixed name, v<V>.metadata.json, and lets whoever succeeds first win. It works where rename is atomic, such as HDFS. But the spec now marks this approach as deprecated and warns that it is not safe on object stores and local file systems, because S3 has no atomic rename.

How it works — a conditional swap

So the standard became the metastore tables approach. The pointer lives in a database or a service and is changed with a check-and-put.

  1. Read the current metadata (version V).
  2. Based on it, write a new metadata <V+1>-<uuid>.metadata.json.
  3. Ask the catalog, "if the pointer is still V, change it to V+1."
  4. If that succeeds, the commit is complete. If it fails, someone else created V+1 first, so start over from step 1.

There are several implementations that honor this contract.

Catalog Where the pointer lives Conditional swap
JDBC One row in the database table iceberg_tables UPDATE only when metadata_location equals the value that was read
REST The catalog server The server checks the requirements and commits
Hive metastore The HMS table property metadata_location Swap together with a metastore lock

The JDBC catalog docs say that the database you connect to must support atomic transactions for atomic commits and serializable reads to be guaranteed. The lab environment of this course keeps this table in a single SQLite file, and Spark (Java's JdbcCatalog) and pyiceberg (SqlCatalog) open the same file. The two engines share one catalog without starting a server.

The REST catalog specification splits a commit into two parts, requirements and updates. The server checks the requirements, such as assert-ref-snapshot-id, first; if they do not hold, it returns a 409 CommitFailedException and the client can retry. Because the server receives the commit, clients need to hold storage credentials less often, and there is also a path for committing several tables atomically at once.

The Hive metastore is easy to attach to the existing Hive ecosystem. As in the Hive docs example, the pointer is in the HMS table property metadata_location. But HMS is a service built for Hive, not for Iceberg, so there are more parts to operate, and the same document says that an INSERT spanning several tables is not atomic.

Names, locations, and files are different layers

One catalog row links only a name and a metadata path. So the following holds.

What it looks like in the field

Moving catalogs. When you move from a Hive metastore to a REST catalog, not one byte of data moves. For each table you read the current metadata path, register it in the new catalog, and drop it from the old catalog without PURGE. The key is to stop the two catalogs from writing to the same table while you move.

A table dropped by mistake. Some people are baffled because space did not shrink after DROP TABLE, and others give up believing everything is gone. If there was no PURGE, a single metadata file path is enough to bring it back.

What really matters in practice

What you will do in the next lab

Spark creates a table in the JDBC catalog (SQLite), pyiceberg opens the same catalog, sees that table, and creates a new table, and then Spark joins the tables of the two engines. You write down how the current and previous paths in one catalog row move before and after a commit, rename a table and confirm that its location stays the same, drop a table without PURGE and count the remaining files, and use register_table to bring a table back from a single metadata file.