TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Calibration Data Decides Accuracy: Narrow, Right, and Wide

Continue in TT Lab

Goal

Implement a CalibrationDataReader yourself, statically quantize the same model with three sets of calibration data — narrow, right and wide — and leave side by side the input scale pinned in the graph and the maximum absolute error you measured. And you confirm by measuring that misaligned calibration cannot be fixed by the number of samples, and that QDQ and QOperator are different shapes of the same computation.

Why it matters

Static quantization pins down the scale of the activations at conversion time, and that scale is decided by the calibration data. So calibration data is an input value, not a setting. Even with the same model and the same command, different calibration data gives a completely different model. Yet that data is left in neither the code nor the logs. There are two directions of collapse. If the range the calibration saw is narrower than the real one, values that went outside stick to the end of the integer data type and do not come back with dequantization. If it is wider, it does not clip, but it divides the grid over a range that is not even used, so the resolution in the interval where values actually live drops. You cannot tell which from a single error; you can tell only by measuring with three sets. And there is a kind that cannot be fixed by the number of samples. If the distribution of the calibration data itself is off, even increasing that data 512 times widens the observed range only a little. Only if you can tell when it is the distribution, not the count, that is the problem, do you not spend time in the wrong place. The grader does not trust the numbers you wrote out. It drives your provider directly to see whether it keeps the contract, tries your tool on a model the grader built, and recomputes the errors you wrote out with your model file to check against them.

Steps

  1. Create and run /root/onnxq-calib/gen_model.py to make /root/onnxq-calib/fp32.onnx.
  2. Create make_reader(arrays, input_name) in /root/onnxq-calib/reader.py.
  3. Create quantize and scales in /root/onnxq-calib/calibtool.py to make /root/onnxq-calib/int8_match.onnx and /root/onnxq-calib/match.json.
  4. Add error and make /root/onnxq-calib/int8_narrow.onnx and /root/onnxq-calib/narrow.json with narrow calibration.
  5. Make /root/onnxq-calib/int8_wide.onnx and /root/onnxq-calib/wide.json with wide calibration.
  6. Gather the three in one place to make /root/onnxq-calib/calib_report.json.
  7. Grow the narrow calibration to 1, 8, 64 and 512 batches and write it in /root/onnxq-calib/samples.json.
  8. Extract the QOperator format from the same calibration too and make /root/onnxq-calib/format.json and /root/onnxq-calib/report.md.

Notes

Get the thing to measure in hand

Create and run /root/onnxq-calib/gen_model.py to make /root/onnxq-calib/fp32.onnx. The batch axis must be a dynamic axis that has only a name, and it is an MLP holding two matrices and two biases.

Build the graph with onnx.helper, check it with onnx.checker and then save. Leave the first axis of the input shape as a name ("N"), not a number. After making it, running it once with onnxruntime makes you sure.

The contract for feeding data into the converter

Create make_reader(arrays, input_name) in /root/onnxq-calib/reader.py. It must return an object that inherits from CalibrationDataReader, and get_next() must give out input dictionaries one at a time and then keep returning None once the data runs out.

The contract is short, but the handling of the end is the key. The converter keeps calling until None comes out, so if you throw an exception after the data runs out, it dies in the middle of calibration. You can hold one iterator and give out with next(it, None), or take items from the front of a list. Remember that an object that has been exhausted once cannot be used again.

Once with the right calibration

Create quantize and scales in /root/onnxq-calib/calibtool.py, make /root/onnxq-calib/int8_match.onnx with seed 4242, scale factor 1.0 and 64 batches, and then write span, batches, input_scale, input_zero_point, observed_span, quantize_nodes and dequantize_nodes in /root/onnxq-calib/match.json.

Give quantize_static quant_format as QDQ and activation_type and weight_type as QInt8. The rule for the calibration data is pinned down in the Notes section — follow it as it is, because the grader has to rebuild the same data. scales finds, in the graph, the QuantizeLinear that takes the input tensor as its first input and reads its scale initializer. Check whether observed_span roughly matches the range width of the calibration data.

Narrow calibration — it gets clipped

Add error, make /root/onnxq-calib/int8_narrow.onnx with seed 4242, scale factor 0.2 and 64 batches, and then write span, batches, input_scale, observed_span, max_abs_error and match_max_abs_error in /root/onnxq-calib/narrow.json. Measure the error with evaluation seed 777 and 256 rows.

You do not multiply the scale factor into the evaluation data — that is the real distribution. If the range the calibration saw is five times narrower than the real one, a good share of the evaluation inputs stick to the end of the data type. If you also write down the error of the right calibration, comparison in the next step is easy.

Wide calibration — it gets squashed

Make /root/onnxq-calib/int8_wide.onnx with seed 4242, scale factor 10.0 and 64 batches, and write /root/onnxq-calib/wide.json in the same shape as narrow.json.

This time there is no clipping — because the calibration saw much wider than the real one. Even so, the error is larger than with the right calibration. It is because the 256 grid steps were distributed over a range that is not even used. Compare observed_span with that of the right calibration.

Put the three in one place

Write eval, narrow, match, wide, best and clipping_worse_than_coarse in /root/onnxq-calib/calib_report.json. The three items must have span, input_scale, observed_span and max_abs_error, and best is the name of the one with the smallest error.

clipping_worse_than_coarse is a boolean holding whether the narrow side's error is larger than the wide side's error. If you set the three errors side by side, you see a U shape — lowest in the middle and rising on both sides. Check which wall is steeper.

What cannot be fixed by the number of samples

Quantize the narrow calibration (scale factor 0.2) four times with 1, 8, 64 and 512 batches and save them as narrow_n1.onnx narrow_n8.onnx narrow_n64.onnx narrow_n512.onnx under /root/onnxq-calib/samples, also save the right calibration with 8 batches as match_n8.onnx, and then write narrow_runs, match_small, narrow_best, narrow_gain and beats_narrow in /root/onnxq-calib/samples.json.

narrow_best is the smallest error of the four runs, and narrow_gain is the 1-batch error divided by narrow_best — that is the improvement gained by increasing the samples 512 times. beats_narrow is a boolean holding whether the right calibration with 8 batches is better than the narrow calibration with 512 batches. 512 batches takes a little time.

The same computation, a different shape

With the right calibration (seed 4242, scale factor 1.0, 64 batches), make the QOperator format /root/onnxq-calib/int8_qop.onnx, write qdq, qoperator, same_input_scale, max_abs_difference and difference_vs_error in /root/onnxq-calib/format.json, and then write /root/onnxq-calib/report.md in four sections.

Put bytes, ops and max_abs_error in the qdq and qoperator items. max_abs_difference is the maximum absolute difference of the outputs when the same evaluation batch is fed into the models of the two formats, and difference_vs_error is that divided by QDQ's error. The error magnitudes of the two formats are of the same order, yet the difference between them is not 0 — because QDQ leaves the way of executing to the runtime. Leave the titles of the four sections of the report exactly as they are in the task sentence.