Count the Float Islands: Operator Support and Substitution
Goal
After exporting the same model five ways, build an analyzer graphscan.py that counts the operators wrapped in int8 and the islands left as float in the graph. You confirm in numbers the effect of the two options that narrow the scope, and fix the model into Gemm to reduce Q/DQ round trips.
Why it matters
A quantization tool converts only the operators it knows into int8 form. In the QDQ format, the operator names remain, and only QuantizeLinear and DequantizeLinear are inserted around tensors. So "quantized" and "runs in int8" are different statements, and to separate the two you must count the graph. This lab fixes the judgment rule as follows. An operator all of whose inputs come from DequantizeLinear and all of whose outputs go only to QuantizeLinear is regarded as wrapped in int8. The rest stay as float, and grouping the connected ones together gives float islands. For islands, the boundary matters more than the count. Each island gets an incoming DQ and an outgoing Q attached, and at each boundary the moving of values and the rounding happen one more time. There are two ways to reduce islands. Narrow the scope and take out the mismatching nodes entirely, or fix the model into a shape that gets quantized. The grader does not trust the numbers you wrote. Each time it builds its own model in a temporary directory with a different number of layers, shape and excluded node, actually runs your analyzer, and checks against the values the grader counted with the same rule. For steps 6 and 7, the grader directly reads the model files you built and counts again.
Steps
- Create and run /root/ops/gen_models.py to make five models under
/root/ops. - Create
nodesin /root/ops/graphscan.py so that it counts the operator census and the initializer data types. - Add
qdqso that it counts QuantizeLinear, DequantizeLinear and Q/DQ round trips. - Add
islandsso that it separates the operators wrapped in int8 from the float islands. - Add
boundaryso that it counts the boundaries going in and out of each island. - Add
diffto output the difference between two models, and write the result of narrowing the scope in /root/ops/scope.json. - Make /root/ops/gemm.onnx, with MatMul and Add merged into Gemm, and /root/ops/gemm_full.onnx, its quantized version, and write them in /root/ops/fuse.json.
- Write /root/ops/ops_report.md in four sections.
Notes
- Python is
/opt/onnx-lab/bin/python. The systempython3has no onnx. - Execution contract:
/opt/onnx-lab/bin/python /root/ops/graphscan.py <명령> <모델...>(the placeholders are the command and the models). On success the exit code is 0, if the file is missing it is 3, and if the usage is wrong it is 2. The answer is output as one JSON blob on standard output. - The five files step 1 makes:
fp32.onnx(the original),full.onnx(static QDQ, the version whose scope is not narrowed),matmul_only.onnx(op_types_to_quantize=["MatMul"]),excluded.onnx(the version with one MatMul taken out withnodes_to_exclude) anddynamic.onnx(quantize_dynamic). The original stacks at least three layers of MatMul, Add and Relu and gives every node a name. The input name isx. nodesresponse:{"model", "nodes", "op_types", "initializers"}. op_types is counts keyed by operator name, and initializers is counts keyed by the data type name of the initializers (names like FLOAT and INT8 given byonnx.TensorProto.DataType.Name).qdqresponse:{"model", "quantize", "dequantize", "round_trips", "weight_dequantize"}. round_trips is the number of QuantizeLinear whose output goes straight into a DequantizeLinear, and weight_dequantize is the number of DequantizeLinear whose first input is an initializer.islandsresponse:{"model", "compute", "wrapped", "float", "island_count", "islands"}. compute is the number of operators excluding QuantizeLinear and DequantizeLinear. Each item of islands is{"nodes", "size", "op_types"}, with nodes and op_types sorted. The island list is sorted by the first node name. A node with an empty name is called연산자이름_자리번호(the placeholders are the operator name and the position number).- Wrapped judgment: a node is wrapped if all of its inputs are outputs of DequantizeLinear and all of its outputs are consumed only by QuantizeLinear. If there is an output with no consumer at all, it is not wrapped.
boundaryresponse:{"model", "island_count", "crossings", "per_island"}. Each item of per_island is{"nodes", "in_dequantize", "out_quantize"}, and the same node is not counted more than once. crossings is the sum of them all.diff <a> <b>response:{"a", "b", "op_delta", "quantize_delta", "dequantize_delta", "round_trip_delta", "island_delta", "float_delta"}. Each delta is b minus a, and op_delta holds only the non-zero items.- The shape of
scope.json:{"excluded_node": 이름, "models": {"full": {...}, "matmul_only": {...}, "excluded": {...}}, "diff_matmul_only": {...}, "diff_excluded": {...}}(where the placeholder is the name). Each item of models is{"quantize","dequantize","round_trips","island_count","float","crossings"}, and the two diffs are thediffresponse as it is, made with a as full.onnx. - The shape of
fuse.json:{"before": {...}, "after": {...}, "removed_quantize": 정수, "removed_round_trips": 정수, "max_abs_error": 실수}(where the placeholders are an integer and a real number). before is for full.onnx and after is for gemm_full.onnx, each{"op_types","quantize","dequantize","round_trips"}. max_abs_error is the maximum absolute error when fp32.onnx and gemm.onnx are run on the same input. - Merging into Gemm: you merge only when the output of
MatMulis used only by thatAddand the other input ofAddis an initializer.Gemm(A, B, C)isA @ B + Cwith the defaults alpha=1.0, beta=1.0, transA=0 and transB=0. After merging, it must passonnx.checker.check_modeland give the same answer as the original. - The dynamic quantization version has a different format.
MatMulIntegerandDynamicQuantizeLinearappear and there is no QuantizeLinear or DequantizeLinear. The yardstick that measures islands and boundaries is for the QDQ version, so for the dynamic version it is honest to use onlynodes. - Official documents: Quantization · Graph optimizations · QuantizeLinear · DequantizeLinear · MatMulInteger · Gemm
- Common mistakes: saying it was quantized from the file size alone, saying it was not quantized because MatMul is visible in the QDQ version, counting the same DequantizeLinear several times, and applying the QDQ yardstick to the dynamic version.
- This lab does not measure time or throughput. It deals only with counting.
Export the same model five ways
Create and run /root/ops/gen_models.py to make fp32.onnx, full.onnx, matmul_only.onnx, excluded.onnx and dynamic.onnx under /root/ops.
Chain MatMul, Add and Relu with onnx.helper and give every node a name. Only with names can you take one out later with nodes_to_exclude. Static quantization needs a CalibrationDataReader, and to use the same calibration data several times, make the reader anew each time — a reader that has been read through once will not hand it out again.
Start by counting what is in it
Create nodes <model> in /root/ops/graphscan.py so that it outputs an operator census and an initializer data type census.
You get the name of an initializer's data type with onnx.TensorProto.DataType.Name(init.data_type). If you count the static QDQ version and the dynamic version side by side, you can see right away how different the formats are — the dynamic version has no MatMul at all.
Count the Q/DQ round trips
Add qdq <model> so that it counts quantize, dequantize, round_trips and weight_dequantize.
A DequantizeLinear attached to a weight has no matching QuantizeLinear. That is because the weights are already in the file as int8. So you must subtract that many from the DequantizeLinear count for it to match the round-trip count of the activations.
Pick out the islands left as float
Add islands <model> so that it separates the operators wrapped in int8 from the operators left as float, and groups the connected ones and outputs them as islands.
If you build a producer and consumer table first, the rest gets easier. Use the rule in the Notes section as it is for the wrapped judgment. When grouping islands, you follow only the connections between float nodes — passing through a Q or DQ means a different island.
Count the boundaries of each island
Add boundary <model> so that it counts the number of DequantizeLinear coming into each island and the number of QuantizeLinear going out, and outputs the total number of boundaries.
The same DequantizeLinear can go into two nodes of one island. Do not count per node; gather and count as a set. If an island's output goes straight out as a graph output, the outgoing QuantizeLinear can be 0.
What changes when you narrow the scope
Add diff <a> <b> so that it outputs the difference between two models, and write the figures of the three versions and the two diffs in /root/ops/scope.json in the shape from the Notes section.
Do not believe it was done just because you gave the option. In the version where you excluded by node name, that node's weights remain as float initializers, and in the version narrowed by operator type, the operators outside that type become an island wholesale. The two kinds of narrowing leave different traces in the graph.
Fix the model to reduce the round trips
Make /root/ops/gemm.onnx, with MatMul and Add merged into a single Gemm, and /root/ops/gemm_full.onnx, which quantizes it with the same settings, and then write the figures of the two versions and the maximum absolute error in /root/ops/fuse.json.
Keep to the merge condition — merge only when the output of MatMul is used only by that Add and the other input of Add is an initializer. After merging, it must pass onnx.checker.check_model, and be sure to run it on the same input as the original and check that the answer is the same. If the values changed, the merge was wrong.
Write it so that the graph can answer
Write /root/ops/ops_report.md in the four sections ## 무엇이 int8 로 갔나 ## float 로 남은 섬 ## 범위를 좁히면 무엇이 달라지나 ## 모델을 고쳐 얻은 것 (the Korean headings mean "What went to int8", "The islands left as float", "What changes when you narrow the scope" and "What you gained by fixing the model"). The number of islands and the number of Q/DQ that disappeared must be in it as numbers.
This report is writing that answers the question "why didn't it shrink as much as expected". If you use the five numbers you counted (operator census, Q/DQ round trips, number of islands, island size, number of boundaries) as they are, that is the answer. Also write that you did not measure time.