Calibration Data Decides Accuracy: Narrow, Right, and Wide
Goal
Implement a CalibrationDataReader yourself, statically quantize the same model with three sets of calibration data — narrow, right and wide — and leave side by side the input scale pinned in the graph and the maximum absolute error you measured. And you confirm by measuring that misaligned calibration cannot be fixed by the number of samples, and that QDQ and QOperator are different shapes of the same computation.
Why it matters
Static quantization pins down the scale of the activations at conversion time, and that scale is decided by the calibration data. So calibration data is an input value, not a setting. Even with the same model and the same command, different calibration data gives a completely different model. Yet that data is left in neither the code nor the logs. There are two directions of collapse. If the range the calibration saw is narrower than the real one, values that went outside stick to the end of the integer data type and do not come back with dequantization. If it is wider, it does not clip, but it divides the grid over a range that is not even used, so the resolution in the interval where values actually live drops. You cannot tell which from a single error; you can tell only by measuring with three sets. And there is a kind that cannot be fixed by the number of samples. If the distribution of the calibration data itself is off, even increasing that data 512 times widens the observed range only a little. Only if you can tell when it is the distribution, not the count, that is the problem, do you not spend time in the wrong place. The grader does not trust the numbers you wrote out. It drives your provider directly to see whether it keeps the contract, tries your tool on a model the grader built, and recomputes the errors you wrote out with your model file to check against them.
Steps
- Create and run /root/onnxq-calib/gen_model.py to make /root/onnxq-calib/fp32.onnx.
- Create
make_reader(arrays, input_name)in /root/onnxq-calib/reader.py. - Create
quantizeandscalesin /root/onnxq-calib/calibtool.py to make /root/onnxq-calib/int8_match.onnx and /root/onnxq-calib/match.json. - Add
errorand make /root/onnxq-calib/int8_narrow.onnx and /root/onnxq-calib/narrow.json with narrow calibration. - Make /root/onnxq-calib/int8_wide.onnx and /root/onnxq-calib/wide.json with wide calibration.
- Gather the three in one place to make /root/onnxq-calib/calib_report.json.
- Grow the narrow calibration to 1, 8, 64 and 512 batches and write it in /root/onnxq-calib/samples.json.
- Extract the QOperator format from the same calibration too and make /root/onnxq-calib/format.json and /root/onnxq-calib/report.md.
Notes
- Python is
/opt/onnx-lab/bin/python. The systempython3has neither numpy nor onnx. Example run:/opt/onnx-lab/bin/python /root/onnxq-calib/calibtool.py scales /root/onnxq-calib/int8_match.onnx. - Execution contract: on success the exit code is 0 and it outputs one JSON blob on standard output. If the number of arguments does not match, it is 2.
quantize <입력.onnx> <출력.onnx> <씨앗> <배율> <묶음수> [형식]response (the placeholders are the input, the output, the seed, the scale factor, the number of batches and the format):{"out", "bytes", "batches", "span", "format", "ops", "input_scale", "input_zero_point"}. The format isqdq(the default) orqoperator. Both activations and weights are QInt8.scales <모델.onnx>response (the placeholder is the model):{"input_scale", "input_zero_point", "observed_span", "quantize_nodes", "dequantize_nodes"}. observed_span is input_scale multiplied by 255.error <fp32.onnx> <int8.onnx> <씨앗> <행>response (the placeholders are the seed and the rows):{"rows", "max_abs_error", "mean_abs_error", "max_abs_output"}.- The rule for making the calibration data (the grader rebuilds the same data): one batch is 16 rows. The number of features is the size of the last axis of the model input. With
rng = numpy.random.default_rng(씨앗)(the placeholder is the seed), you make(rng.normal(0, 1, (16, 특징)) * 배율).astype(numpy.float32)(where the placeholders are the features and the scale factor) for each batch and produce as many as the number of batches. You cast down to float32 after multiplying. - The rule for making the evaluation data: it is
numpy.random.default_rng(씨앗).normal(0, 1, (행, 특징)).astype(numpy.float32)(where the placeholders are the seed, the rows and the features). You do not multiply by the scale factor — this is the real distribution. - The values this lab uses: calibration seed 4242, evaluation seed 777, 256 evaluation rows, and 64 batches. The scale factors are 0.2 for narrow, 1.0 for right and 10.0 for wide.
make_reader(arrays, input_name)must return an object that inherits fromonnxruntime.quantization.CalibrationDataReader.get_next()returns a{입력이름: 배열}dictionary (the placeholders are the input name and the array) one at a time and returnsNonewhen the data runs out. It must keep returning None even if called again after the end — if it throws an exception, the converter dies in the middle of calibration.- A provider that has been exhausted cannot be reused. Make a new one for each quantization.
- The input scale pinned in the model can be read by finding, in the graph, the
QuantizeLinear(orQLinearMatMulfor the QOperator format) that takes the input tensor as its first input and reading its scale initializer. calibtool.pyimportsreader.pyfrom the same directory. Put the directory where the script is intosys.path.- The models that step 7 makes are saved under /root/onnxq-calib/samples as
narrow_n1.onnxnarrow_n8.onnxnarrow_n64.onnxnarrow_n512.onnxmatch_n8.onnx. 512 batches takes a little time. - Calling
quantize_staticproduces a warning recommending preprocessing. The model in this lab is already simple, so it works as it is without preprocessing. - Official documents: ONNX Runtime quantization · QuantizeLinear · DequantizeLinear · ONNX Concepts
- Common mistakes: making a provider once and reusing it across several conversions, throwing an exception after the end, multiplying the scale factor into the evaluation data, and guessing that different formats will have different accuracy.
Get the thing to measure in hand
Create and run /root/onnxq-calib/gen_model.py to make /root/onnxq-calib/fp32.onnx. The batch axis must be a dynamic axis that has only a name, and it is an MLP holding two matrices and two biases.
Build the graph with onnx.helper, check it with onnx.checker and then save. Leave the first axis of the input shape as a name ("N"), not a number. After making it, running it once with onnxruntime makes you sure.
The contract for feeding data into the converter
Create make_reader(arrays, input_name) in /root/onnxq-calib/reader.py. It must return an object that inherits from CalibrationDataReader, and get_next() must give out input dictionaries one at a time and then keep returning None once the data runs out.
The contract is short, but the handling of the end is the key. The converter keeps calling until None comes out, so if you throw an exception after the data runs out, it dies in the middle of calibration. You can hold one iterator and give out with next(it, None), or take items from the front of a list. Remember that an object that has been exhausted once cannot be used again.
Once with the right calibration
Create quantize and scales in /root/onnxq-calib/calibtool.py, make /root/onnxq-calib/int8_match.onnx with seed 4242, scale factor 1.0 and 64 batches, and then write span, batches, input_scale, input_zero_point, observed_span, quantize_nodes and dequantize_nodes in /root/onnxq-calib/match.json.
Give quantize_static quant_format as QDQ and activation_type and weight_type as QInt8. The rule for the calibration data is pinned down in the Notes section — follow it as it is, because the grader has to rebuild the same data. scales finds, in the graph, the QuantizeLinear that takes the input tensor as its first input and reads its scale initializer. Check whether observed_span roughly matches the range width of the calibration data.
Narrow calibration — it gets clipped
Add error, make /root/onnxq-calib/int8_narrow.onnx with seed 4242, scale factor 0.2 and 64 batches, and then write span, batches, input_scale, observed_span, max_abs_error and match_max_abs_error in /root/onnxq-calib/narrow.json. Measure the error with evaluation seed 777 and 256 rows.
You do not multiply the scale factor into the evaluation data — that is the real distribution. If the range the calibration saw is five times narrower than the real one, a good share of the evaluation inputs stick to the end of the data type. If you also write down the error of the right calibration, comparison in the next step is easy.
Wide calibration — it gets squashed
Make /root/onnxq-calib/int8_wide.onnx with seed 4242, scale factor 10.0 and 64 batches, and write /root/onnxq-calib/wide.json in the same shape as narrow.json.
This time there is no clipping — because the calibration saw much wider than the real one. Even so, the error is larger than with the right calibration. It is because the 256 grid steps were distributed over a range that is not even used. Compare observed_span with that of the right calibration.
Put the three in one place
Write eval, narrow, match, wide, best and clipping_worse_than_coarse in /root/onnxq-calib/calib_report.json. The three items must have span, input_scale, observed_span and max_abs_error, and best is the name of the one with the smallest error.
clipping_worse_than_coarse is a boolean holding whether the narrow side's error is larger than the wide side's error. If you set the three errors side by side, you see a U shape — lowest in the middle and rising on both sides. Check which wall is steeper.
What cannot be fixed by the number of samples
Quantize the narrow calibration (scale factor 0.2) four times with 1, 8, 64 and 512 batches and save them as narrow_n1.onnx narrow_n8.onnx narrow_n64.onnx narrow_n512.onnx under /root/onnxq-calib/samples, also save the right calibration with 8 batches as match_n8.onnx, and then write narrow_runs, match_small, narrow_best, narrow_gain and beats_narrow in /root/onnxq-calib/samples.json.
narrow_best is the smallest error of the four runs, and narrow_gain is the 1-batch error divided by narrow_best — that is the improvement gained by increasing the samples 512 times. beats_narrow is a boolean holding whether the right calibration with 8 batches is better than the narrow calibration with 512 batches. 512 batches takes a little time.
The same computation, a different shape
With the right calibration (seed 4242, scale factor 1.0, 64 batches), make the QOperator format /root/onnxq-calib/int8_qop.onnx, write qdq, qoperator, same_input_scale, max_abs_difference and difference_vs_error in /root/onnxq-calib/format.json, and then write /root/onnxq-calib/report.md in four sections.
Put bytes, ops and max_abs_error in the qdq and qoperator items. max_abs_difference is the maximum absolute difference of the outputs when the same evaluation batch is fed into the models of the two formats, and difference_vs_error is that divided by QDQ's error. The error magnitudes of the two formats are of the same order, yet the difference between them is not 0 — because QDQ leaves the way of executing to the runtime. Leave the titles of the four sections of the report exactly as they are in the task sentence.