Schemas Change: Surviving Whichever Side Ships First
One-line summary
The real question in a schema change is not "what do we change" but "does it not break whichever of the writer and the reader changes first," and the answer is decided by three things: defaults, aliases, and promotion rules.
Why this was needed
You need to add one more field to the events a pipeline emits. You fix the schema file and deploy the producer. That night three downstream systems stop. One of them is a team our team did not even know existed.
Next time you are careful: you fix the reader first and deploy the writer later. This time the reader stops because it cannot find the new field. This is because nobody is using that field yet.
The two incidents have the same cause. A deployment does not happen in an instant. Between the writer and the reader there is always a period when only one side has the new code, and during that period data keeps flowing. So what you should ask when designing a change is not "is this change correct" but "does it survive being deployed in either order."
This is where you must separate this from the earlier module of this course. Detecting that a file someone else gives you changed without notice is the receiving side's defense. Here we are the side making the change, and it is about keeping old data and old reading code alive at the same time while we number our own versions and change them.
How it works
We call the two directions by name. We use the names that Confluent's compatibility documentation uses, as they are.
- Backward compatible: reading code that uses the new schema reads data written with the old schema. If this works, you may upgrade the reader first.
- Forward compatible: reading code that uses the old schema reads data written with the new schema. If this works, you may upgrade the writer first.
- If both work, it is fully compatible, and only then can you deploy the two sides independently.
The judgment rules are already written in the schema resolution section of the Avro specification. Three lines are all there is.
- A field that exists only on the writer side is simply dropped by the reader. So adding a field does not break old reading code.
- A field that exists only on the reader side must have a default. If it does not, it is an error. So if you add a required field without a default, new reading code cannot read old data.
- The type of the same field must be promoted from the writer to the reader. The promotions the specification allows are int to long, float, and double, long to float and double, and float to double. The opposite direction is not a promotion.
Practical rules follow directly from these three lines. Adding a field with a default is safe in both directions. When the reader meets old data, it fills in with the default, and old reading code only needs to drop the new field. Conversely, adding a required field without a default breaks backward compatibility. This is because old data does not have that value at all and the reader has no way to fill it in.
Renaming is more subtle. Looking at the schema alone, a rename is one addition and one deletion. The only device that turns it back into a single event is the alias. But Avro uses aliases only from the reader's schema — because it works by rewriting the writer's schema into the reader's names when reading. So a rename survives in only one direction. New reading code holds the alias, so it reads old data, but old reading code has no alias pointing to the new name. The reason Protocol Buffers uses field numbers instead of names and states that a number, once used, cannot be changed is the same.
Type changes also split by direction. If you widen int to long, backward compatibility holds and forward compatibility breaks. If you narrow long to int, it is exactly the opposite. So the phrase "we changed the type" decides nothing by itself, and you have to look at which direction it was changed.
What it looks like in the field
First, writing code that reads only one version. During the transition period, several versions flow mixed together. If the reader knows only the latest version, it throws away the old data wholesale for that whole period, and the fact of having thrown it away shows up only as a reduced count. There is only one way to get through the transition period — attach generous defaults to the read schema so that even the old versions can be read.
Second, not writing the version number in the data. If there is no mark on each line of which version it was written with, the reader can only guess, and guesses are wrong. The version number must travel with the data.
Third, deleting a field without looking at its default. If you delete a field, old reading code cannot find that field. If that field had a default in the old schema, it gets filled in, and otherwise it breaks. So the only fields you can delete are those that had a default from the start, and this fact is decided in advance when you add the field.
Fourth, managing compatibility only in documents. A document does not stop a deployment. You must turn the judgment into code and make that code run first when you release a new version. If you give the reason as a fixed code rather than a sentence for people to read, automation becomes easy.
What really matters in practice
- Start by choosing to add a field with a default. It is the only change that is safe in both directions.
- Do not rename, and if you must, decide the aliases and the deployment order together. The reader goes first.
- Use only the widening direction for types. Narrowing holds only for old reading code, and usually the opposite is needed.
- Write the version number on every line, and cover several versions with one read schema. The transition period is not an exception but the normal state.
What to do in the next lab
After creating and emitting five versions of order events, you build up the contract tool contract.py step by step. You separate required from optional by default, classify the changes between versions as additions, deletions, renames, and type changes, judge backward and forward compatibility according to the Avro rules, and produce the reasons as fixed codes. At the end, you read lines from the five mixed versions with a strict read schema and a tolerant read schema, and see in numbers how many rows one default saves. The grader builds its own schemas each time with different field names and types, actually runs your tool, and compares the answers.