TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

No Calibration Data: What Dynamic Quantization Does to the Graph

Continue in TT Lab

In one line

Dynamic quantization fixes only the weights to int8 in advance and measures the range of the activations on the spot at every inference. So you do not need calibration data, and in return you use a different ruler for each batch.

Why this branch was needed

To do static quantization, you need representative inputs. That is because you have to look in advance at roughly what range the activations live in and pin down the scale. But getting that data takes the longest in the field. You cannot get customer data, synthetic data has a different distribution, and the data you do get is mixed with the evaluation data.

Dynamic quantization sidesteps that problem entirely. It does not decide the range of the activations in advance but looks at that tensor at execution time and decides. All the conversion needs is one model file, and one command line finishes it. So "let's first see how far this gets us" becomes possible, and it serves as the baseline for deciding whether to do static quantization at all.

How it works

The ONNX Runtime quantization documentation explains the two branches side by side. On the dynamic side, only the weights are converted to integers offline, and the scale and zero point of the activations are computed during execution.

If you open the graph, you see right away what happened. A single MatMul changes like this.

바뀌기 전   X(float32) ──▶ MatMul(W float32) ──▶ h(float32)

바뀐 뒤     X(float32) ──▶ DynamicQuantizeLinear ──▶ (q uint8, scale, zero_point)
                                 │
                                 ▼
                           MatMulInteger(W int8) ──▶ (int32)
                                 │
                            Cast ──▶ Mul(scale 들) ──▶ h(float32)

MatMul disappears and DynamicQuantizeLinear, MatMulInteger, Cast and Mul come in its place. The first node brings the activation down to uint8 on the spot while producing the scale and zero point, the integer kernel multiplies, and then it is multiplied by the two scales to come back to a real number.

The definition of DynamicQuantizeLinear is short. It sets the range by adding 0 to the minimum and maximum of the data, divides that width by 255 to make the scale, and takes the integer position where 0.0 will land as the zero point. That this computation happens again at every inference is the whole of this method.

Where the file size shrinks is also fixed. If you count the initializers one by one, only the matrices shrink and the biases stay as float32. And next to the shrunken matrix, the scale and zero point follow as small initializers. The common belief that "quantizing makes it a quarter" is true only when the matrices are most of the file.

What it looks like in the field

First, it holds up even when the input size changes. Even if you feed in an input 1000 times larger, the relative error stays nearly the same. That is because it remakes the ruler for every batch. Static quantization in the same situation goes outside the range seen during calibration and gets clipped wholesale. The biggest advantage of this method is that there is no risk that the calibration data differs from the real distribution.

Second, neighbors mixed into the same batch ruin each other. The scale is decided by the maximum of that whole tensor, so if even one row with an exceptionally large value comes in, the grid the other rows can use shrinks. An input that was fine when fed alone gets a worse answer once it is in a batch — which means the result differs depending on the batch composition for the same input, and where reproducibility is needed, this immediately becomes a problem.

Third, in a small model the file actually gets bigger. If you apply dynamic quantization to a model whose weights are only a few hundred bytes, the newly added nodes and the scale and zero point initializers are larger than the amount that shrank. If you move on thinking "it must have gotten smaller" without measuring, it gets deployed as it is.

Fourth, measuring the range happens at every inference. To find the scale, you have to sweep the activation tensor from start to end to find the minimum and maximum. The number of DynamicQuantizeLinear in the graph is exactly how many times that sweep runs per inference. Counting the nodes first before measuring time tells you where the cost attaches.

What really matters in practice

What you will do in the next lab

You build dyntool.py to dynamically quantize your own model, count the graph change and the bytes per initializer, implement by hand the computation DynamicQuantizeLinear does at every inference, measure the error while shaking the input scale by 1000 times, measure how one outlier row shakes the batch, and confirm that the file gets bigger in a very small model. The grader builds the model and batches anew each time with a different seed and shape, actually runs your tool, and redoes the same computation to check against it.