TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

No Calibration Data: Measuring Dynamic Quantization

Continue in TT Lab

Goal

Build dyntool.py to apply dynamic quantization to your own model, and leave as numbers how the graph changes, where the file shrinks and what stays the same, why it holds up even when the input distribution changes, and what the case where this method does not fit looks like.

Why it matters

Static quantization can begin only after you gather representative inputs, and getting that data takes the longest in the field. Dynamic quantization does not decide the range of the activations in advance but looks at that tensor and decides at every inference, so all you need is one model file. So it becomes the baseline you try first before deciding what to do. In exchange, you have to know what you gain and what you lose in order to choose. What you gain is the property of holding up even when the distribution shakes — it remakes the ruler for every batch, so it never goes outside the range seen during calibration. What you lose is reproducibility. The scale is decided by the maximum of that whole batch, so one outlier row pulls down even the other rows in the same batch. An input that was fine when fed alone gets worse because of its neighbor. You also have to measure the file size to know. What shrinks is only the matrices, the biases stay float32, and next to the shrunken matrix the scale and zero point follow. In a model with small weights, the newly added nodes are larger than the amount that shrank, so the file actually gets bigger. The grader does not trust the numbers you wrote out. It builds the model anew each time with a different seed and shape, actually runs your tool, and checks against the actual result of the DynamicQuantizeLinear operator and the inference result the grader ran itself.

Steps

  1. Create and run /root/onnxq-dyn/gen_model.py to make /root/onnxq-dyn/fp32.onnx.
  2. Create quantize in /root/onnxq-dyn/dyntool.py to make /root/onnxq-dyn/dyn.onnx.
  3. Add sizes to count the bytes per initializer and write it in /root/onnxq-dyn/size.json.
  4. Add dqparams and implement by hand the computation that DynamicQuantizeLinear does at every inference.
  5. Add compare to measure the error while changing the input scale factor and write it in /root/onnxq-dyn/batches.json.
  6. Add outlier to measure whether one outlier row shakes the batch and write it in /root/onnxq-dyn/outlier.json.
  7. Build a very small model with /root/onnxq-dyn/gen_small.py, quantize it and write it in /root/onnxq-dyn/unfit.json.
  8. Summarize on one page with /root/onnxq-dyn/report.json and /root/onnxq-dyn/report.md.

Notes

Get the thing to measure in hand

Create and run /root/onnxq-dyn/gen_model.py to make /root/onnxq-dyn/fp32.onnx. The batch axis must be a dynamic axis that has only a name, and it is an MLP holding two matrices and two biases.

Build the graph with onnx.helper, check it with onnx.checker and then save. Leave the first axis of the input shape as a name ("N"), not a number. Make the initializers with numpy_helper.from_array. After making it, running it once with onnxruntime makes you sure.

Fix only the weights with one command line

Create quantize <입력.onnx> <출력.onnx> (the placeholders are the input and the output) in /root/onnxq-dyn/dyntool.py to make /root/onnxq-dyn/dyn.onnx. The response must have the operator counts after the change, the operators that disappeared, and the operators that newly came in.

Just give quantize_dynamic in onnxruntime.quantization a weight_type of QInt8. You do not put in calibration data — there is no place for it. You can see the changed graph by opening it with onnx.load and counting the op_type of the nodes. Check where MatMul went.

Say by name where it shrank

Add sizes <모델.onnx> (the placeholder is the model) so that it counts the data type, number of elements and bytes for each initializer, and write the result of comparing the two models in /root/onnxq-dyn/size.json as fp32_bytes, dyn_bytes, shrunk, kept_float, weight_bytes_before and weight_bytes_after.

If you go through graph.initializer of onnx.load and convert with numpy_helper.to_array, you get dtype and nbytes. shrunk is the initializers that exist in fp32 but whose names are gone in the dynamic model, and kept_float is those that stayed as float32 as they were. See which side the biases are on.

The computation that remakes the ruler at every inference

Add dqparams <입력.npy> <출력.npy> (the placeholders are the input and the output) and implement by hand the computation that DynamicQuantizeLinear does at every inference. It outputs the scale, zero point and range and saves the quantized uint8 array.

You must always put 0 in the range — if the minimum is positive, bring it down to 0, and if the maximum is negative, bring it up to 0. It is uint8, so there are 255 grid steps, and the zero point is the integer position where the real number 0.0 lands. If you put in an all-positive array and an all-negative array, you can see the zero point going to either end.

Shake the input by 1000 times

Add compare <fp32.onnx> <int8.onnx> <씨앗> <행> <배율> (the placeholders are the seed, the rows and the scale factor), measure scale factors 0.05, 1.0 and 50.0 with seed 20260917 and 32 rows, and write it in /root/onnxq-dyn/batches.json as runs and relative_spread.

The rule for building the batch is pinned down in the Notes section — follow it as it is, because the grader has to rebuild the same batch. relative_spread is the maximum of the three relative values divided by the minimum. The absolute error grows with the input size, but see what the relative error does.

One neighbor row shakes the batch

Add outlier <fp32.onnx> <int8.onnx> <씨앗> <행> <튀는값> (the placeholders are the seed, the rows and the spike value), measure with seed 20260917, 12 rows and spike value 60.0, and write it in /root/onnxq-dyn/outlier.json.

You feed in the same 12 rows twice — once as they are, and once with a spike row attached below. You measure the error only on the first 12 rows. You do not count the spike row's own error. Only then does the statement 'it got worse because of its neighbor' hold.

In a small model it actually grows

Build a very small model /root/onnxq-dyn/small.onnx with /root/onnxq-dyn/gen_small.py, quantize it to make /root/onnxq-dyn/small_dyn.onnx, and then write the result in /root/onnxq-dyn/unfit.json as small_fp32_bytes, small_dyn_bytes, grew, growth_bytes, dynamic_quantize_nodes and kept_float_ops.

If the weights are only a few hundred bytes, the newly added nodes and the scale and zero point initializers are larger than the amount that shrank. kept_float_ops are the operators that are in both the fp32 model and the dynamic model — those that remained unchanged. grew is a boolean, not a string.

What was gained and what was lost, on one page

Write graph_change, file, relative_error_spread, outlier_ratio and small_model_grew in /root/onnxq-dyn/report.json, and write /root/onnxq-dyn/report.md in the four sections ## 그래프가 어떻게 바뀌나 ## 어디가 줄고 어디가 안 줄었나 ## 보정 자료가 필요 없는 이유 ## 동적 양자화가 맞지 않는 경우 (the Korean headings mean "How the graph changes", "Where it shrank and where it did not", "Why calibration data is not needed" and "Where dynamic quantization does not fit").

You can assemble it by reading the JSON files you made in the earlier steps. The ratio of file is the dynamic model bytes divided by the fp32 bytes. In the report, write along with each number one line on what that number refers to — the reader has not seen your experiment.