TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Scale and Zero Point: Quantization Is Choosing a Ruler

Continue in TT Lab

In one line

Quantization is deciding the grid spacing (scale) and the position of 0 (zero_point) on the real-number axis, and once those two are decided, the rest is only division and rounding.

Why you need to know this

Call a conversion tool and an int8 model comes out. When accuracy drops, what should you fix? Give less calibration data, split by channel, leave only the activations as float? These choices are all choices about how the scale is decided. If you do not know what arithmetic runs inside, you cannot choose and only trial and error remains.

And there is a place where it really goes wrong. The most common is the case where 0 does not come back as 0. If you set the scale using the observed range [0.4, 7.0] as it is, the real number 0.0 is outside the representable grid. The 0 that padding and ReLU produced comes back as a non-zero value, and that error swells as it goes through the layers. That is why an implementation that follows the specification always includes 0 in the range.

How it works

ONNX does this work with two operators. QuantizeLinear brings real numbers down to integers, and DequantizeLinear raises integers to real numbers. The definitions are short.

QuantizeLinear     y = saturate(round(x / y_scale) + y_zero_point)
DequantizeLinear   x = (q - x_zero_point) * x_scale

round is rounding that sticks to the even side. 0.5 goes to 0, 1.5 goes to 2, and 2.5 also goes to 2. saturate is cutting to the two ends of the data type, which is 0 and 255 for uint8 and -128 and 127 for int8.

The scale and zero point come from the observed range [lo, hi]. There are two branches.

Two rulers that fold the real-number axis onto an integer grid. Asymmetric uint8 distributes codes 0 through 255 over the observed range from minus 0.5 to 2.0, and the position where the real number 0 lands becomes the zero point 51. Symmetric int8 pins 0 in the middle so the zero point is 0, and if the data is only 0 or above, the left half of 127 steps sits idle

An important property follows here. The round-trip error does not exceed half the scale. Quantization moves a value to the nearest point on a grid of multiples of the scale, and since the grid spacing is the scale, at most it is half. But that holds only when it is within the range. A value that went outside sticks to the wall, so the error can be arbitrarily large. If the error greatly exceeds half the scale, it is a range problem, not a rounding problem.

What it looks like in the field

First, handling activations symmetrically throws away half. The activations after a ReLU are all 0 or above. If you apply symmetric int8 here, the 127 grid steps on the negative side receive no values. The remaining grid is half, so the scale becomes twice as coarse, and the error becomes exactly double. That is why in many cases the default is symmetric for weights and asymmetric for activations.

Second, one outlier ruins the whole tensor. If only one position in a weight matrix has a large value, the scale of the whole tensor grows to fit that value, and all the other positions crowd onto just a few grid steps. Here the answer is per-channel scale. If you give each output channel its own scale, the damage is confined within the channel that holds the outlier. The axis attribute of QuantizeLinear and a 1-dimensional scale tensor express this.

Third, if you compute the scale in float64, it goes off from the runtime. The runtime divides in float32. If the value divided with a Python float lands on a boundary, the rounding goes the other way, and about 1 percent of the elements differ by one grid step. It is not a big problem, but it is a place where you lose time asking "why does my calculation differ from the model output".

Fourth, to measure error you need a criterion. "Accuracy dropped by 1 percent" is a deployment judgment, and "the round-trip error exceeded half the scale" is a cause judgment. The latter can be known by looking at a single layer and points directly at where to fix.

What really matters in practice

What you will do in the next lab

You grow qmath.py one step at a time and implement scale and zero point computation, quantization, dequantization, round-trip error and per-channel scale yourself. The grader does not trust the numbers you wrote out — each time it builds arrays with a different seed, actually runs your tool, and checks against the result of putting the same values straight through ONNX's QuantizeLinear and DequantizeLinear operators.