No Calibration Data: Measuring Dynamic Quantization
Goal
Build dyntool.py to apply dynamic quantization to your own model, and leave as numbers how the graph changes, where the file shrinks and what stays the same, why it holds up even when the input distribution changes, and what the case where this method does not fit looks like.
Why it matters
Static quantization can begin only after you gather representative inputs, and getting that data takes the longest in the field. Dynamic quantization does not decide the range of the activations in advance but looks at that tensor and decides at every inference, so all you need is one model file. So it becomes the baseline you try first before deciding what to do. In exchange, you have to know what you gain and what you lose in order to choose. What you gain is the property of holding up even when the distribution shakes — it remakes the ruler for every batch, so it never goes outside the range seen during calibration. What you lose is reproducibility. The scale is decided by the maximum of that whole batch, so one outlier row pulls down even the other rows in the same batch. An input that was fine when fed alone gets worse because of its neighbor. You also have to measure the file size to know. What shrinks is only the matrices, the biases stay float32, and next to the shrunken matrix the scale and zero point follow. In a model with small weights, the newly added nodes are larger than the amount that shrank, so the file actually gets bigger. The grader does not trust the numbers you wrote out. It builds the model anew each time with a different seed and shape, actually runs your tool, and checks against the actual result of the DynamicQuantizeLinear operator and the inference result the grader ran itself.
Steps
- Create and run /root/onnxq-dyn/gen_model.py to make /root/onnxq-dyn/fp32.onnx.
- Create
quantizein /root/onnxq-dyn/dyntool.py to make /root/onnxq-dyn/dyn.onnx. - Add
sizesto count the bytes per initializer and write it in /root/onnxq-dyn/size.json. - Add
dqparamsand implement by hand the computation that DynamicQuantizeLinear does at every inference. - Add
compareto measure the error while changing the input scale factor and write it in /root/onnxq-dyn/batches.json. - Add
outlierto measure whether one outlier row shakes the batch and write it in /root/onnxq-dyn/outlier.json. - Build a very small model with /root/onnxq-dyn/gen_small.py, quantize it and write it in /root/onnxq-dyn/unfit.json.
- Summarize on one page with /root/onnxq-dyn/report.json and /root/onnxq-dyn/report.md.
Notes
- Python is
/opt/onnx-lab/bin/python. The systempython3has neither numpy nor onnx. Example run:/opt/onnx-lab/bin/python /root/onnxq-dyn/dyntool.py sizes /root/onnxq-dyn/fp32.onnx. - Execution contract: on success the exit code is 0 and it outputs one JSON blob on standard output. If the number of arguments does not match, it is 2.
quantize <입력.onnx> <출력.onnx>response (the placeholders are the input and the output):{"out": 경로, "bytes": 정수, "ops": {연산자: 개수}, "removed": [사라진 연산자], "added": [새로 든 연산자]}(where the placeholders are a path, an integer, an operator and a count, the disappeared operators, and the newly added operators). The weights are QInt8.sizes <모델.onnx>response (the placeholder is the model):{"file_bytes": 정수, "initializers": {이름: {"dtype": 문자열, "elements": 정수, "bytes": 정수}}, "initializer_bytes": 정수}(where the placeholders are integers, a name and a string).dqparams <입력.npy> <출력.npy>response (the placeholders are the input and the output):{"scale": 실수, "zero_point": 정수, "min": 실수, "max": 실수, "count": 정수}(where the placeholders are real numbers and integers). It saves the quantized uint8 array to the output path.compare <fp32.onnx> <int8.onnx> <씨앗> <행> <배율>response (the placeholders are the seed, the rows and the scale factor):{"rows", "span", "max_abs_output", "max_abs_error", "relative"}. relative is the error divided by the maximum absolute value of the fp32 output.outlier <fp32.onnx> <int8.onnx> <씨앗> <행> <튀는값>response (the placeholders are the seed, the rows and the spike value):{"rows", "spike", "clean", "with_spike", "ratio"}.- The rule for building the batch (the grader rebuilds the same batch): the number of features is the size of the last axis of the model input. It is
rng = numpy.random.default_rng(씨앗)(the placeholder is the seed) with(rng.normal(0, 1, (행, 특징)) * 배율).astype(numpy.float32)(where the placeholders are the rows, the features and the scale factor). You cast down to float32 after multiplying. - The rule of
outlier: build the batch using the rule above with a scale factor of 1.0, and clean is the maximum absolute error when only that batch is fed in. with_spike is the maximum absolute error measured on the first as many rows as the row count, after feeding in the batch with one row ofnumpy.full((1, 특징), 튀는값, dtype=numpy.float32)(where the placeholders are the features and the spike value) attached below. ratio is with_spike divided by clean. - The definition of DynamicQuantizeLinear: the range is
[min(x, 0), max(x, 0)], the scale is that width divided by 255, and the zero point is the integer position where the real number 0.0 lands, clipped to between 0 and 255. Rounding is the same round-half-to-even as numpy'snp.rint. - Make the scale with
np.float32(...)and do the division in float32 too. If you divide with a Python float, elements on a boundary go off by one grid step from the runtime. - The batch axis of
fp32.onnxmust be a dynamic axis that has only a name (["N", 20]). If you fix it to a number, the grader cannot feed in other numbers of rows. - Step 5 measures the scale factors 0.05, 1.0 and 50.0 with seed 20260917 and 32 rows. Step 6 uses seed 20260917, 12 rows and spike value 60.0.
- Official documents: ONNX Runtime quantization · DynamicQuantizeLinear · MatMulInteger · ONNX Concepts
- Calling
quantize_dynamicproduces a warning recommending preprocessing. The model in this lab is already simple, so it works as it is without preprocessing. - Common mistakes: guessing the file shrank and not measuring, writing that the biases also became int8, swapping the order of casting down to float32 in the batch rule, and in the outlier experiment measuring the error including the spike row itself.
Get the thing to measure in hand
Create and run /root/onnxq-dyn/gen_model.py to make /root/onnxq-dyn/fp32.onnx. The batch axis must be a dynamic axis that has only a name, and it is an MLP holding two matrices and two biases.
Build the graph with onnx.helper, check it with onnx.checker and then save. Leave the first axis of the input shape as a name ("N"), not a number. Make the initializers with numpy_helper.from_array. After making it, running it once with onnxruntime makes you sure.
Fix only the weights with one command line
Create quantize <입력.onnx> <출력.onnx> (the placeholders are the input and the output) in /root/onnxq-dyn/dyntool.py to make /root/onnxq-dyn/dyn.onnx. The response must have the operator counts after the change, the operators that disappeared, and the operators that newly came in.
Just give quantize_dynamic in onnxruntime.quantization a weight_type of QInt8. You do not put in calibration data — there is no place for it. You can see the changed graph by opening it with onnx.load and counting the op_type of the nodes. Check where MatMul went.
Say by name where it shrank
Add sizes <모델.onnx> (the placeholder is the model) so that it counts the data type, number of elements and bytes for each initializer, and write the result of comparing the two models in /root/onnxq-dyn/size.json as fp32_bytes, dyn_bytes, shrunk, kept_float, weight_bytes_before and weight_bytes_after.
If you go through graph.initializer of onnx.load and convert with numpy_helper.to_array, you get dtype and nbytes. shrunk is the initializers that exist in fp32 but whose names are gone in the dynamic model, and kept_float is those that stayed as float32 as they were. See which side the biases are on.
The computation that remakes the ruler at every inference
Add dqparams <입력.npy> <출력.npy> (the placeholders are the input and the output) and implement by hand the computation that DynamicQuantizeLinear does at every inference. It outputs the scale, zero point and range and saves the quantized uint8 array.
You must always put 0 in the range — if the minimum is positive, bring it down to 0, and if the maximum is negative, bring it up to 0. It is uint8, so there are 255 grid steps, and the zero point is the integer position where the real number 0.0 lands. If you put in an all-positive array and an all-negative array, you can see the zero point going to either end.
Shake the input by 1000 times
Add compare <fp32.onnx> <int8.onnx> <씨앗> <행> <배율> (the placeholders are the seed, the rows and the scale factor), measure scale factors 0.05, 1.0 and 50.0 with seed 20260917 and 32 rows, and write it in /root/onnxq-dyn/batches.json as runs and relative_spread.
The rule for building the batch is pinned down in the Notes section — follow it as it is, because the grader has to rebuild the same batch. relative_spread is the maximum of the three relative values divided by the minimum. The absolute error grows with the input size, but see what the relative error does.
One neighbor row shakes the batch
Add outlier <fp32.onnx> <int8.onnx> <씨앗> <행> <튀는값> (the placeholders are the seed, the rows and the spike value), measure with seed 20260917, 12 rows and spike value 60.0, and write it in /root/onnxq-dyn/outlier.json.
You feed in the same 12 rows twice — once as they are, and once with a spike row attached below. You measure the error only on the first 12 rows. You do not count the spike row's own error. Only then does the statement 'it got worse because of its neighbor' hold.
In a small model it actually grows
Build a very small model /root/onnxq-dyn/small.onnx with /root/onnxq-dyn/gen_small.py, quantize it to make /root/onnxq-dyn/small_dyn.onnx, and then write the result in /root/onnxq-dyn/unfit.json as small_fp32_bytes, small_dyn_bytes, grew, growth_bytes, dynamic_quantize_nodes and kept_float_ops.
If the weights are only a few hundred bytes, the newly added nodes and the scale and zero point initializers are larger than the amount that shrank. kept_float_ops are the operators that are in both the fp32 model and the dynamic model — those that remained unchanged. grew is a boolean, not a string.
What was gained and what was lost, on one page
Write graph_change, file, relative_error_spread, outlier_ratio and small_model_grew in /root/onnxq-dyn/report.json, and write /root/onnxq-dyn/report.md in the four sections ## 그래프가 어떻게 바뀌나 ## 어디가 줄고 어디가 안 줄었나 ## 보정 자료가 필요 없는 이유 ## 동적 양자화가 맞지 않는 경우 (the Korean headings mean "How the graph changes", "Where it shrank and where it did not", "Why calibration data is not needed" and "Where dynamic quantization does not fit").
You can assemble it by reading the JSON files you made in the earlier steps. The ratio of file is the dynamic model bytes divided by the fp32 bytes. In the report, write along with each number one line on what that number refers to — the reader has not seen your experiment.