TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Count the Float Islands: Operator Support and Substitution

Continue in TT Lab

Goal

After exporting the same model five ways, build an analyzer graphscan.py that counts the operators wrapped in int8 and the islands left as float in the graph. You confirm in numbers the effect of the two options that narrow the scope, and fix the model into Gemm to reduce Q/DQ round trips.

Why it matters

A quantization tool converts only the operators it knows into int8 form. In the QDQ format, the operator names remain, and only QuantizeLinear and DequantizeLinear are inserted around tensors. So "quantized" and "runs in int8" are different statements, and to separate the two you must count the graph. This lab fixes the judgment rule as follows. An operator all of whose inputs come from DequantizeLinear and all of whose outputs go only to QuantizeLinear is regarded as wrapped in int8. The rest stay as float, and grouping the connected ones together gives float islands. For islands, the boundary matters more than the count. Each island gets an incoming DQ and an outgoing Q attached, and at each boundary the moving of values and the rounding happen one more time. There are two ways to reduce islands. Narrow the scope and take out the mismatching nodes entirely, or fix the model into a shape that gets quantized. The grader does not trust the numbers you wrote. Each time it builds its own model in a temporary directory with a different number of layers, shape and excluded node, actually runs your analyzer, and checks against the values the grader counted with the same rule. For steps 6 and 7, the grader directly reads the model files you built and counts again.

Steps

  1. Create and run /root/ops/gen_models.py to make five models under /root/ops.
  2. Create nodes in /root/ops/graphscan.py so that it counts the operator census and the initializer data types.
  3. Add qdq so that it counts QuantizeLinear, DequantizeLinear and Q/DQ round trips.
  4. Add islands so that it separates the operators wrapped in int8 from the float islands.
  5. Add boundary so that it counts the boundaries going in and out of each island.
  6. Add diff to output the difference between two models, and write the result of narrowing the scope in /root/ops/scope.json.
  7. Make /root/ops/gemm.onnx, with MatMul and Add merged into Gemm, and /root/ops/gemm_full.onnx, its quantized version, and write them in /root/ops/fuse.json.
  8. Write /root/ops/ops_report.md in four sections.

Notes

Export the same model five ways

Create and run /root/ops/gen_models.py to make fp32.onnx, full.onnx, matmul_only.onnx, excluded.onnx and dynamic.onnx under /root/ops.

Chain MatMul, Add and Relu with onnx.helper and give every node a name. Only with names can you take one out later with nodes_to_exclude. Static quantization needs a CalibrationDataReader, and to use the same calibration data several times, make the reader anew each time — a reader that has been read through once will not hand it out again.

Start by counting what is in it

Create nodes <model> in /root/ops/graphscan.py so that it outputs an operator census and an initializer data type census.

You get the name of an initializer's data type with onnx.TensorProto.DataType.Name(init.data_type). If you count the static QDQ version and the dynamic version side by side, you can see right away how different the formats are — the dynamic version has no MatMul at all.

Count the Q/DQ round trips

Add qdq <model> so that it counts quantize, dequantize, round_trips and weight_dequantize.

A DequantizeLinear attached to a weight has no matching QuantizeLinear. That is because the weights are already in the file as int8. So you must subtract that many from the DequantizeLinear count for it to match the round-trip count of the activations.

Pick out the islands left as float

Add islands <model> so that it separates the operators wrapped in int8 from the operators left as float, and groups the connected ones and outputs them as islands.

If you build a producer and consumer table first, the rest gets easier. Use the rule in the Notes section as it is for the wrapped judgment. When grouping islands, you follow only the connections between float nodes — passing through a Q or DQ means a different island.

Count the boundaries of each island

Add boundary <model> so that it counts the number of DequantizeLinear coming into each island and the number of QuantizeLinear going out, and outputs the total number of boundaries.

The same DequantizeLinear can go into two nodes of one island. Do not count per node; gather and count as a set. If an island's output goes straight out as a graph output, the outgoing QuantizeLinear can be 0.

What changes when you narrow the scope

Add diff <a> <b> so that it outputs the difference between two models, and write the figures of the three versions and the two diffs in /root/ops/scope.json in the shape from the Notes section.

Do not believe it was done just because you gave the option. In the version where you excluded by node name, that node's weights remain as float initializers, and in the version narrowed by operator type, the operators outside that type become an island wholesale. The two kinds of narrowing leave different traces in the graph.

Fix the model to reduce the round trips

Make /root/ops/gemm.onnx, with MatMul and Add merged into a single Gemm, and /root/ops/gemm_full.onnx, which quantizes it with the same settings, and then write the figures of the two versions and the maximum absolute error in /root/ops/fuse.json.

Keep to the merge condition — merge only when the output of MatMul is used only by that Add and the other input of Add is an initializer. After merging, it must pass onnx.checker.check_model, and be sure to run it on the same input as the original and check that the answer is the same. If the values changed, the merge was wrong.

Write it so that the graph can answer

Write /root/ops/ops_report.md in the four sections ## 무엇이 int8 로 갔나 ## float 로 남은 섬 ## 범위를 좁히면 무엇이 달라지나 ## 모델을 고쳐 얻은 것 (the Korean headings mean "What went to int8", "The islands left as float", "What changes when you narrow the scope" and "What you gained by fixing the model"). The number of islands and the number of Q/DQ that disappeared must be in it as numbers.

This report is writing that answers the question "why didn't it shrink as much as expected". If you use the five numbers you counted (operator census, Q/DQ round trips, number of islands, island size, number of boundaries) as they are, that is the answer. Also write that you did not measure time.