TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Not Every Operator Makes It to int8

Continue in TT Lab

In one line

A quantization tool converts only the operators it knows into int8 form. The rest stay as float, and at each such boundary a Q/DQ round trip arises.

Why this becomes a problem

After running quantize_static, if the file size has shrunk, it feels as if the job is done. But if you open the graph, half of it is as before. A quantization tool has a table that lists which operators it handles and how, and operators not in that table are left untouched. An operator left untouched must be computed in float, so the values made int8 before it are unpacked back to float, and then packed into int8 again after it.

So "quantized" and "runs in int8" are different statements. If you believe they are the same, it does not get as fast or as small as you expected and you cannot find the reason. The reason is all written inside the graph — you just need to count.

How to read a QDQ version

The QDQ format of static quantization does not change the operators. Instead, it inserts QuantizeLinear and DequantizeLinear at each tensor. The graph still shows MatMul written as float, and the executor recognizes the bundle DQ -> 연산자 -> Q (the Korean word means operator) and swaps it for an int8 kernel and runs it.

So to judge "what can go to int8" from looking at the graph alone, you need a rule. This lab decides it this way. An operator all of whose inputs come from DequantizeLinear and all of whose outputs go only to QuantizeLinear is regarded as wrapped in int8. If even one thing is off, that operator stays as float.

감싸인 모양                        남은 모양
  DQ -> MatMul -> Q                 DQ -> MatMul -> Add -> Relu -> Q
        (int8 커널로 접힌다)                    (Add·Relu 는 float 섬)

If the operators left as float are connected to one another, they are one lump. Let us call this a float island. For islands, the count matters, but the number of boundaries matters more. Each island gets an incoming DequantizeLinear and an outgoing QuantizeLinear attached. As islands increase, boundaries increase, and at each boundary the moving of values and the rounding happen one more time.

One more thing stands out in actual measurement. If you quantize without narrowing the scope, Relu disappears from the graph. It is not surprising; it has been folded into a boundary. If you put the zero point of the following QuantizeLinear at -128, the lower end of int8, the range of values it can hold starts exactly from 0, and negatives are clipped to 0 the moment they are stored. The boundary is doing what Relu used to do. If you measure it on this lab's model, the zero point of the hidden layer is actually -128 and the representable range starts from 0.0. If you do not count, you cannot explain why Relu is 0 in the census.

One more. Dynamic quantization is a completely different shape. MatMul changes to MatMulInteger and DynamicQuantizeLinear is attached. If you apply the yardstick you used to measure the QDQ version to this as it is, you get a wrong answer. The yardstick must be set anew for each format.

Three things you do in the field

First, narrow the scope. If accuracy collapses at some layer, you just take out that node. The tool takes node names with nodes_to_exclude and operator types with op_types_to_quantize. If you exclude by name, that node's weights stay as float, and if you narrow by type, the operators outside that type become a float island wholesale. Either way, you must count to confirm what changed — do not believe it was done just because you gave the option.

Second, fix the model into a shape that gets quantized. The shape where Add comes after MatMul is very common, and if you merge the two into a single Gemm, the intermediate tensor that was between them disappears. When the intermediate tensor disappears, the pair of Q/DQ attached to that tensor disappears too. The values stay the same and only the boundaries decrease. In actual measurement, on a 3-layer model, QuantizeLinear went from 7 to 4.

There is something people often miss when narrowing the scope. Narrowing is not free. If you take out a node, that node runs in float, but the Q/DQ before and after it mostly remain. That is, you get a spot where you lose the gain of int8 and pay the cost of the boundary. So the judgment "accuracy wobbles, so let's take it out for now" must also be made after counting. If you put the number of islands and boundaries after taking it out side by side with the un-narrowed version, what you lost and what you gained is left as numbers.

Third, report by counting. The operator census, the number of Q/DQ and round trips, the number and size of islands, and the number of boundaries. With these five numbers you can answer the question "why didn't it shrink as much as expected" from the graph. Reading the graph first before measuring time is much faster.

What really matters in practice

What you will do in the next lab

After exporting the same model five ways, you grow step by step an analyzer that counts the graph. You count in turn the operator census, the Q/DQ round trips, the float islands and the number of boundaries per island, and produce the difference between the two versions with narrowed scope and the original. At the end you merge MatMul and Add into Gemm, quantize again, and count how many Q/DQ disappeared. The grader builds its own model each time with a different number of layers and a different excluded node, actually runs your analyzer, and counts the same numbers to check against them.