TT Lab
Get started
Learn Learning paths Courses

Data Pipelines

Schemas Change: Surviving Whichever Side Ships First

Continue in TT Lab

One-line summary

The real question in a schema change is not "what do we change" but "does it not break whichever of the writer and the reader changes first," and the answer is decided by three things: defaults, aliases, and promotion rules.

Why this was needed

You need to add one more field to the events a pipeline emits. You fix the schema file and deploy the producer. That night three downstream systems stop. One of them is a team our team did not even know existed.

Next time you are careful: you fix the reader first and deploy the writer later. This time the reader stops because it cannot find the new field. This is because nobody is using that field yet.

The two incidents have the same cause. A deployment does not happen in an instant. Between the writer and the reader there is always a period when only one side has the new code, and during that period data keeps flowing. So what you should ask when designing a change is not "is this change correct" but "does it survive being deployed in either order."

This is where you must separate this from the earlier module of this course. Detecting that a file someone else gives you changed without notice is the receiving side's defense. Here we are the side making the change, and it is about keeping old data and old reading code alive at the same time while we number our own versions and change them.

How it works

We call the two directions by name. We use the names that Confluent's compatibility documentation uses, as they are.

The judgment rules are already written in the schema resolution section of the Avro specification. Three lines are all there is.

Practical rules follow directly from these three lines. Adding a field with a default is safe in both directions. When the reader meets old data, it fills in with the default, and old reading code only needs to drop the new field. Conversely, adding a required field without a default breaks backward compatibility. This is because old data does not have that value at all and the reader has no way to fill it in.

Renaming is more subtle. Looking at the schema alone, a rename is one addition and one deletion. The only device that turns it back into a single event is the alias. But Avro uses aliases only from the reader's schema — because it works by rewriting the writer's schema into the reader's names when reading. So a rename survives in only one direction. New reading code holds the alias, so it reads old data, but old reading code has no alias pointing to the new name. The reason Protocol Buffers uses field numbers instead of names and states that a number, once used, cannot be changed is the same.

Type changes also split by direction. If you widen int to long, backward compatibility holds and forward compatibility breaks. If you narrow long to int, it is exactly the opposite. So the phrase "we changed the type" decides nothing by itself, and you have to look at which direction it was changed.

What it looks like in the field

First, writing code that reads only one version. During the transition period, several versions flow mixed together. If the reader knows only the latest version, it throws away the old data wholesale for that whole period, and the fact of having thrown it away shows up only as a reduced count. There is only one way to get through the transition period — attach generous defaults to the read schema so that even the old versions can be read.

Second, not writing the version number in the data. If there is no mark on each line of which version it was written with, the reader can only guess, and guesses are wrong. The version number must travel with the data.

Third, deleting a field without looking at its default. If you delete a field, old reading code cannot find that field. If that field had a default in the old schema, it gets filled in, and otherwise it breaks. So the only fields you can delete are those that had a default from the start, and this fact is decided in advance when you add the field.

Fourth, managing compatibility only in documents. A document does not stop a deployment. You must turn the judgment into code and make that code run first when you release a new version. If you give the reason as a fixed code rather than a sentence for people to read, automation becomes easy.

What really matters in practice

What to do in the next lab

After creating and emitting five versions of order events, you build up the contract tool contract.py step by step. You separate required from optional by default, classify the changes between versions as additions, deletions, renames, and type changes, judge backward and forward compatibility according to the Avro rules, and produce the reasons as fixed codes. At the end, you read lines from the five mixed versions with a strict read schema and a tolerant read schema, and see in numbers how many rows one default saves. The grader builds its own schemas each time with different field names and types, actually runs your tool, and compares the answers.