TT Lab
Get started
Learn Learning paths Courses

Transformers — Compute Attention By Hand

What Changes When You Fold Into int8 and Back

Continue in TT Lab

In one line

Quantization is turning a real number into one scale and one integer, and what you lose is decided by the scale. What decides that scale is the single largest absolute value in the list.

Why this was needed

If you convert a model to int8, the weights shrink to a quarter and you can use integer multipliers. So everyone tries it once. The problem is what comes after — accuracy drops a little, and you cannot say where it dropped.

The tool ends in one line. You call a quantization function and a model comes out. If you do not know what arithmetic happened inside, all you can do is change options and run it again, and that is not fixing but testing your luck.

That is why this text and the next lab do not use a tool. They start by folding and unfolding one list of eight values by hand. What you see there shows up just the same in large models.

What folding means

To move a list of real numbers into int8, you have to decide two things. The scale and the zero point.

The simplest is symmetric quantization. You set the scale to max(|x|) / 127, divide the values by the scale and round. The largest absolute value reaches code 127, and 0 lands exactly on code 0.

scale = max(abs(x) for x in values) / 127
codes = [round(x / scale) for x in values]   # -127 부터 127 까지
back  = [c * scale for c in codes]           # 편 값. 원본이 아니다

Asymmetric (affine) quantization is used when the values are skewed to one side. You set the width to (max - min) / 255 so that you use all 256 slots, and carry separately, as the zero point, which code the real number 0 lands on. Restoring is (code - zero_point) * scale. The paper that lays out inference with integer arithmetic only uses exactly this formula.

In both methods, what you lose is in the same place. Half of the scale. If the scale is 0.007, any value is off by at most 0.0035. Whether the value is large or small, it is exactly that much.

First pin down the rounding

This is where people get caught first. Python's round is not the rounding you learned in school.

round(0.5)   # 0   — 1 이 아니다
round(1.5)   # 2
round(2.5)   # 2   — 3 이 아니다

It sends the position at exactly 0.5 to the even side. It is defined this way because if you always round 0.5 up, the average of the rounded values is pushed up little by little. Quantization is rounding every element of an array once, so this push becomes the model's bias as it is.

The problem is not the rule but the fact that there are two rules. Some implementations send to the even side and some send to the side away from 0. If you fold the same weights with the same scale and the codes differ by one slot, every later comparison loses its meaning. That is why, when reading quantization code, the rounding rule is the first thing to check, before the scale formula.

What decides the scale is one value

The scale of symmetric quantization is max(|x|) / 127. This formula has neither a mean nor a variance. One largest absolute value is everything.

So if one exceptionally large value is mixed into the list, that one decides the precision of all the rest. If the rest are all below 1 and one is 42, the scale becomes 42/127, and the values below 1 get squashed into codes 0 to 3. Even with 256 slots, only four are used.

A figure in which a single 42 among eight values decides the scale. If you use one scale for the whole matrix, the seven values below 1 get squashed into four slots, codes 0 to 3, and if you set the scale separately for each row, the row without the 42 uses the scale properly, dividing it from code 28 to 127

This is the starting point of the LLM.int8() paper. It was the observation that in the hidden states of large language models, components far larger than the other values appear, and because of those few, everything else becomes unusable. The direction of the solution also comes from there — set the scale in smaller units, or pull the large ones out separately.

Setting the scale in smaller units is what to do first. Instead of folding the whole matrix with one scale, if you set the scale separately for each row (or each column), a small row is not dragged along by a large row. But it is not free. You have to carry a scale for every row, and for a matrix multiplication to hold, the scale must be per row or per column. If each value has a different scale, you cannot pull the scales out from inside the integer accumulation.

What remains when you multiply with integers

The heart of integer inference is that multiplication and accumulation are all integer. While you multiply and add codes, rounding never happens even once. After adding everything up, at the end you multiply by the two scales and unfold to a real number.

acc  = sum(a_code[k] * b_code[k] for k in range(d))   # 정수만
value = acc * a_scale * b_scale                       # 마지막에 한 번

So the error that arises in a matrix multiplication is not something that swelled up in the accumulation but something that already arose when folding at the start. That means there is only one place to look for the cause, and this is good news.

The remaining question is what happens to that error later. The attention scores pass through the softmax and become probabilities. The softmax spreads differences out exponentially, so a small error in the scores can grow in the probabilities, or it can instead get buried. Which one it is cannot be said before measuring. You measure it yourself in the next lab.

What it looks like in the field

First, accuracy dropped a little and you do not know where it dropped. The tool is one line, so you have no eyes to look inside. Only someone who has done by hand what moves when you change the scale, the rounding or the unit can point to it.

Second, you quantized the same model with two tools and the results differ. Even if the scale formula is the same, if the rounding rule differs, the codes are off by one slot. You should first check what differs, not which side is right.

Third, you folded at the tensor level and it collapses only in certain layers. This is when the weight distribution of that layer has an exceptionally large value. The overall mean error looks fine, but only the relative error of the small rows blows up.

Fourth, activations are much trickier than weights. Weights are fixed, so you measure once and you are done, but activations differ for each input. You measure the range with calibration data, and if that data differs from real input, the scale goes off.

Fifth, people talk about size and speed as the same thing. The file becoming a quarter and it actually getting faster are different matters. Whether the integer kernel actually ran has to be checked separately.

What really matters in practice

What you will do in the next lab

You grow /root/work/tf-quant/quant.py one step at a time. You do not call a tool and do the arithmetic yourself — the system Python of this Pod has no numpy (numpy exists only inside /opt/onnx-lab). So you handle lists and lists of lists with the standard library. floor and exp of Python's math module are enough.

You start by building the two rounding rules and confirming where they part ways. Then you build symmetric and asymmetric quantization, and build a yardstick that measures how far the folded-and-unfolded values have moved from the original. That is half of it.

The other half is the point of this lab. You attach one large value to ordinary values and count into how many slots the other values shrink. With a matrix whose rows differ greatly in width, you measure per-tensor and per-row units side by side. You multiply codes with integers only and confirm that the accumulation is exactly right, and finally you measure what that error becomes in the probabilities after passing through the softmax. You measure once more with a key that has an outlier planted in it, and see how what you saw earlier shows up in attention.

The grader does not trust the explanations you wrote down. It actually imports your module, pokes at the functions with different inputs every time, and checks them against values it computes separately. Most of the checks are between integer arrays — if even one code slot is off, it shows immediately.