TT Lab
Get started
Learn Learning paths Courses

I renumbered a field and the old client silently read the wrong value

Contract checks belong to a script, not a reviewer

Continue in TT Lab

Summary

A diff where one number changed does not catch the human eye. That is why schema compatibility checks must be done by a script, not by a reviewer — parse two .proto files, compare the numbers per message, and transcribe the document's list of "unsafe changes" into a form a machine can read.

Why this was needed

In the previous module we reproduced the incident in the title. A change that renumbers compiles, passes unit tests, and in code review looks like a line where two numbers changed. A reviewer cannot remember every time "what was this number before?". Even less does anyone remember the number of a field deleted half a year ago. When the best practices document writes "never reuse a number, even if you think nobody uses it", it means that such memory is not trustworthy.

Machines remember. If you put the old .proto and the new .proto side by side and compare the numbers, you catch everything that people miss. This check is a transcription of the document's message type update rules, so there is no room for judgment, and if you put it in CI it runs before merging. In this module you build that checker yourself with the standard library.

How it works

What to compare. Wire compatibility is viewed per message, with the field number as the key. The name is secondary. So the checker's data structure is sufficient as {메시지: {번호: (이름, 타입)}} (message → number → (name, type)) plus the sets of reserved numbers and reserved names. Nested messages are named Outer.Inner and compared separately — if you look only at the outer message, you miss deletions in the inner one.

The rules come from the document. The table below is the verdict the checker emits, and the right column is the basis.

Verdict Condition Basis
REMOVED_NOT_RESERVED The old number is not in the new file and is not reserved either Deleting is safe but the number must not be reused → block it with reserved
RENUMBERED The same name under a different number Changing a number is the same as deleting and creating anew — it is unsafe
TYPE_CHANGED The same number and the same name, a different type Almost never change a type (some are conditionally compatible)
NUMBER_REUSED The same number with a different name and a different type Reusing a number makes decoding ambiguous
REUSED_RESERVED The new file uses as a field a reserved number of the old file Do not take a number out of the reserved list and use it
WARN RENAMED The same number and the same type, only the name differs There is no name on the wire — binaries are safe, JSON and TextProto are affected

The reason to treat a rename only as a warning is in the encoding document — the bytes have only numbers and wire types, and the reader attaches names by looking at the .proto. If you change only the name, not a single bit of the bytes changes. However, it is a breaking change for consumers using formats that serialize names, such as ProtoJSON, so the checker reports it but does not block. If you cannot make this distinction, the checker raises a red light every time, and a checker with frequent red lights gets ignored.

Rule precedence. One cause should yield one line. If id moves from 1 to 2, "1 disappeared (REMOVED)" and "2 received a different name (RENAMED)" are simultaneously true too, but the cause is one — the number was changed — so it emits only RENUMBERED and does not look at that number further. If the number remains in the new file, it is not REMOVED.

Lenient parser, strict verdict. Real .proto files mix comments, declarations spanning several lines, options such as [deprecated = true], and ranges such as reserved 9 to 11. If the parser stumbles on these, people turn the checker off. The verdict, on the other hand, yields not an inch — it reads even if the format is slightly different, but it always fails number reuse. In Python, you can erase comments with a regex, find message 이름 { (message name) and match the braces by counting them, and read the inside with a 타입 이름 = 번호; (type name = number;) regex.

FIELD = re.compile(r"(?:\b(optional|repeated)\s+)?([A-Za-z_][\w.]*)\s+([A-Za-z_]\w*)\s*=\s*(\d+)\s*(?:\[[^\]]*\])?\s*;")

Output contract. The checker is a tool. It produces both lines for people to read and an exit code for machines to read — a violation is one per line, starting with the verdict name and carrying the message name and number. If there are violations, exit 1; if none, print OK on the last line and exit 0. Warnings do not affect the exit code. With this contract, you can sweep a directory with a shell script and block CI if even one fails.

What the checker cannot catch. The document's conditionally compatible list — changing int32 to int64 — is safe only when you control the deployment order. The checker does not know that order, so it is right to block it as TYPE_CHANGED. Conversely, a change in the meaning of a default value ("0 now means 'free' rather than 'undecided'") changes neither the bytes nor the schema text, so no checker can catch it. This is why the best practices document separately writes "almost never change a field's default value". The checker does not replace review; it keeps the reviewer from spending time comparing numbers.

What you meet in the field

There is something you usually experience the day you first put a schema compatibility check into CI — dozens of old schemas in the repository turn red at once. This is because deleted numbers without reservation have piled up over years. If you turn the rule off at this point, it becomes the same as having no tool. Instead, you first submit a cleanup PR that fills reserved into the old files. Reserving deleted numbers is a change that is safe at any time, so there is no risk.

The second is "warning fatigue". If you treat a rename as an error, every time someone fixes a typo the check is blocked, and people learn --no-verify. After that, real violations slip through too. Separating warnings from errors is not a luxury but a condition for the checker to survive.

The third is nested messages. The checker was looking only at top-level messages, so a type change of Order.Item.count was merged as is. The one line that names things Outer.Inner when you write the parser prevents this incident.

What you will do in the next lab

You build /root/grpc/compat/protocheck.py. First, with --dump, you check that the parser reads the fixtures and a file it has never seen (comments, range reservations, nesting, options), and then add the rules one by one — deletion without reservation, renumbering, type change and number reuse, reserved number reuse, and renaming as a warning only. You run the 8 fixture pairs to make a report, and finally build check-all.sh, which runs per directory. The grader also makes .proto pairs that are not in the instructions, besides the fixtures, and feeds them to your script, so you cannot hard-code fixture names.