TT Lab
Get started
Learn Learning paths Courses

I renumbered a field and the old client silently read the wrong value

One line of bytes carries the field number

Continue in TT Lab

Summary

Protobuf bytes contain no field names. The single tag byte 08 means "field 1, varint", and the 96 01 after it is 150. The names and declared types are attached by the reader's .proto — so if you change a number, the parser reads another field's value without any error.

Why this was needed

JSON sends the name along, like {"qty": 3}. It is easy for people to read, and renaming a field is noticed immediately. In exchange, every message has to carry the name as a string, and numbers also have to be written as strings. For internal services with heavy traffic, this cost is visible.

Protocol Buffers made the opposite choice. It does not send names but only numbers. It also sends values not as text but in the shortest binary representation. The price is that the reader must know the schema (.proto), and the numbers in the schema become the contract. This whole course is about that price. If you read the bytes by hand once, you understand why in your bones.

This lab image has neither protoc nor the protobuf Python package. That is actually a good thing. The rules set by the encoding document can all be transcribed in a hundred lines of the Python standard library, and the purpose of this module is to open up by hand what the library hides.

How it works

Let us follow the document's first example exactly. If you serialize message Test1 { int32 a = 1; } with a = 150, three bytes come out.

08 96 01

Start with the varint. You cut the integer into 7-bit groups and emit them from the lowest, and if there is more after, you set the most significant bit (MSB) of that byte. 150 in binary is 10010110. Putting the continuation mark on the lower 7 bits 0010110 gives 10010110 = 0x96, and the remaining 1 is 0x01. So it is 96 01. 1 is the single byte 01, 300 is ac 02, and 16384 is 80 80 01 — the smaller the value, the shorter. The values below were computed directly in Python while writing this text.

    1 → 01
  127 → 7f
  128 → 80 01
  150 → 96 01
  300 → ac 02
16384 → 80 80 01

The tag is (field_number << 3) | wire_type encoded as a varint. The lower 3 bits are the wire type and the rest is the field number. 08 is 0000 1000, so wire type 0, number 1. There are six wire types, and you actually meet four of them.

Number Name Types that use it
0 VARINT int32, int64, uint32, uint64, sint32, sint64, bool, enum
1 I64 fixed64, sfixed64, double
2 LEN string, bytes, submessages, packed repeated
5 I32 fixed32, sfixed32, float

(3 SGROUP and 4 EGROUP are for the deprecated groups.) For field numbers 1 through 15 the tag is one byte, and from 16 on it becomes two bytes — (16 << 3) | 0 = 128, so it is 80 01. This calculation is why the proto3 language guide says "give 1 through 15 to frequently used fields".

What the wire type does is tell you the length. VARINT runs up to the byte whose MSB is 0, I64 is 8 bytes, I32 is 4 bytes, and LEN is as many as the varint right after it. Thanks to this, even when an unknown number arrives, the parser knows how much to skip. This is the principle by which an old binary passes over a new field without error, and we will measure that property in the next module.

A LEN record is the tag, then a length varint, then the payload. For message Test2 { string b = 2; } with "testing" put in, it looks like this.

12 07 74 65 73 74 69 6e 67
│  │  └── "testing" 의 UTF-8 7바이트
│  └───── 길이 7
└──────── 태그 (2 << 3) | 2 = 0x12

The length is the number of bytes, not the number of characters. A five-character Korean memo is 13 bytes in UTF-8, so the length varint is also 13. If you count it as the number of characters, the parser reads the next tag from the wrong place.

A submessage is also LEN. In message Test3 { Test1 c = 3; }, if c.a = 150, it is 1a 03 08 96 01 — the tag 1a (3:LEN), length 3, and then those same three bytes from before sit inside as they are. If you decode the payload once more with the same function, the inner message comes out.

Negative numbers are the trap. int32 and int64 treat negative numbers as two's complement, and the document explains this as "as an unsigned 64-bit integer the most significant bit is set, so it uses all ten bytes". -1 is ff ff ff ff ff ff ff ff ff 01, ten bytes. For fields where negatives are frequent, use sint32. sint converts first with ZigZag — a positive p becomes 2p and a negative n becomes 2|n|-1. It alternates 0→0, -1→1, 1→2, -2→3, and the 32-bit formula is (n << 1) ^ (n >> 31). Then -1 becomes varint 1, that is, the single byte 01. -150 is ZigZag 299 → the two bytes ab 02. You will measure directly in the lab how the same value splits into 10 bytes and 2 bytes.

A missing field simply has no record written. There is no header and no field count. A message is nothing but records concatenated, so field order is not guaranteed either and the parser must read them in any order. The document states firmly "do not assume that the serialized bytes are stable" — serializing the same message twice is not guaranteed to give the same bytes, so you must not put a hash or CRC on top of it.

Packed repeated is the default from Edition 2023. Instead of attaching a tag to every element, a repeated scalar field concatenates the values into a single LEN record. If you put 1, 2, 3 into repeated int32 e = 5, it is 2a 03 01 02 03 — three values under one tag. The document requires the parser to read both the packed form and the per-element form.

What you meet in the field

The most common incident is "the data is corrupted but there is no error". A team renumbered the fields to make them look nice while tidying up the schema. It compiled and the tests passed. But a consumer service still using the old binary began to read the order number in the place of the order quantity. From the parser's point of view, a varint arrived for number 2, so it simply read it by the rules. Names are not on the wire, so there is no information at all that could say "this is not qty". An incident like this leaves no exception in the logs, so it takes days to discover.

The second is a size misjudgment. We calculated capacity assuming "4 bytes per integer", but negative int32 values went out at ten bytes each and the result came to more than double the estimate. Conversely, after switching to sint32 it sometimes shrinks to less than half. The habit of looking at the distribution of values when choosing a type comes from here.

The third is string length. The first bug anyone who writes a parser by hand encounters is almost without exception a confusion between UTF-8 byte count and character count. It never shows up if you test only ASCII, and it blows up on the first Korean memo.

What you will do in the next lab

In /root/grpc/pb.py, you build in turn a varint, tag, ZigZag, field, and message encoder and decoder. You measure directly that 150 becomes 96 01, and that int32 -1 is ten bytes and sint32 -1 is one byte. Then you decode fixture bytes to pick out the fields the v1 schema does not know, and decode a submessage and a packed repeated field twice. The grader imports your pb.py and round-trips it with random values not in the instructions, so you cannot hard-code values. You have to transcribe the rules.