Build the Ruler Yourself: Scale, Zero Point, Round-Trip Error
Goal
Build qmath.py to compute the scale and zero point exactly as the definitions of QuantizeLinear and DequantizeLinear, implement quantization and dequantization, confirm that the round-trip error is within half the scale, and compare symmetric and asymmetric, and per-channel scale and whole-tensor scale, on the same data.
Why it matters
To pick where to fix when the accuracy of the int8 model a conversion tool produced drops, you have to know the arithmetic the tool does inside. The scale is the observed range divided by the number of grid steps, and the zero point is the integer position where the real number 0.0 lands. If 0 is not in the range, the 0 that padding and ReLU made comes back as a non-zero value.
The round-trip error has an upper bound. As long as the value is within the range, the error does not exceed half the scale — because quantization moves a value to a grid of the scale spacing. So if the error greatly exceeds that half, it is the range, not the rounding, that is the problem. This one yardstick makes the cause judgment fast.
The shape of the data also decides the method. The activations after a ReLU are all 0 or above, so if you apply symmetric int8, the grid on the negative side sits idle as a whole and the scale becomes twice as coarse. If there is one outlier in a weight matrix, the whole-tensor scale gets dragged to that value and everything else is squashed — if you split the channels, the damage is confined within that channel.
The grader does not trust the numbers you wrote out. Each time it builds arrays with a different seed, actually runs your qmath.py, and checks against the result of putting the same input straight through ONNX's QuantizeLinear and DequantizeLinear operators.
Steps
- Create and run /root/onnxq-scale/gen_data.py to make three sets of arrays under /root/onnxq-scale/data, and implement
paramsin /root/onnxq-scale/qmath.py. - Add
qto bring a real-number array down to an integer array. - Add
dqto turn an integer array back into real numbers. - Add
roundtripso that it outputs the round-trip error and its upper bound together. - Deliberately narrow the range to create saturation and write the result in /root/onnxq-scale/saturate.json.
- Apply symmetric and asymmetric to data that is all 0 or above and write it in /root/onnxq-scale/symmetry.json.
- Add
channelto compare the per-channel scale and the whole-tensor scale and write it in /root/onnxq-scale/channel.json. - Summarize on one page with /root/onnxq-scale/report.json and /root/onnxq-scale/report.md.
Notes
- Python is
/opt/onnx-lab/bin/python. The systempython3has no numpy. Example run:/opt/onnx-lab/bin/python /root/onnxq-scale/qmath.py params -3.2 5.1 asym-u8. - Execution contract: on success the exit code is 0 and it outputs one JSON blob on standard output. If the number of arguments does not match, it is 2.
params <lo> <hi> <mode>response:{"scale": 실수, "zero_point": 정수, "qmin": 정수, "qmax": 정수, "mode": 문자열}(where the placeholders are a real number, an integer and a string).q <입력.npy> <출력.npy> <mode> [lo hi]response (the placeholders are the input and the output):{"scale": 실수, "zero_point": 정수, "clipped": 정수, "dtype": 문자열, "count": 정수}(where the placeholders are a real number, integers and a string). If you do not give lo and hi, it uses the minimum and maximum of the input array. clipped is the number of elements that were cut.dq <양자.npy> <출력.npy> <scale> <zero_point> <mode>response (the placeholders are the quantized array and the output):{"min": 실수, "max": 실수, "count": 정수}(where the placeholders are real numbers and an integer).roundtrip <입력.npy> <mode> [lo hi]response (the placeholder is the input):{"mode", "scale", "zero_point", "lo", "hi", "clipped", "max_abs_error", "bound", "within_bound"}. bound is half the scale and within_bound is a boolean.channel <입력.npy> <axis>response (the placeholder is the input):{"axis", "channels", "per_tensor_scale", "per_tensor_max_error", "per_channel_scales", "per_channel_max_error", "worst_channel", "channels_improved"}. It computes with symmetric int8.- There are two modes.
asym-u8is uint8, qmin 0, qmax 255, asymmetric, andsym-i8is int8, qmin -128, qmax 127, symmetric with zero point 0. In symmetric, you divide one side's width by 127 (not 128). - Always include 0 in the range. If lo is greater than 0, bring it down to 0, and if hi is less than 0, raise it to 0.
- The rounding is rounding that sticks to the even side. numpy's
np.rintbehaves that way. If you use theround()builtin orint(x + 0.5), the answers differ. - When making the scale, wrap it with
np.float32(...)and do the division in float32 too. If you divide with a Python float, elements on a boundary go off by one grid step from the runtime. - The material arrays are three files under /root/onnxq-scale/data: spread.npy, positive.npy and weights.npy. gen_data.py makes them, and the grader reads these files.
- Official documents: QuantizeLinear · DequantizeLinear · ONNX Concepts · ONNX Runtime quantization
- Common mistakes: not putting 0 in the range, dividing by 128 in symmetric, using round-half-down rounding instead of
np.rint, and writing that it does not exceed the bound even when saturation occurred.
Turn the observed range into a scale and zero point
Create and run /root/onnxq-scale/gen_data.py to make three arrays under /root/onnxq-scale/data, and implement params <lo> <hi> <mode> in /root/onnxq-scale/qmath.py. If 0 is not in the range, you must put it in.
For asymmetric, divide (hi - lo) by 255, and the zero point is the integer position where the real number 0.0 lands — round qmin minus lo/scale. For symmetric, divide the larger of the left and right widths by 127 and the zero point is 0. Handle the cases where lo is greater than 0 or hi is less than 0 first.
Bring real numbers down to an integer grid
Add q <입력.npy> <출력.npy> <mode> [lo hi] (the placeholders are the input and the output) so that it brings a real-number array down to an integer array and saves the result as .npy. The response must have scale, zero_point, clipped, dtype and count.
The definition is saturate(round(x / scale) + zero_point). Divide, round, add the zero point, and cut to the two ends of the data type. If you swap the order, the answer changes — you must not add the zero point first and then round. Count the number of elements that were cut and give it as clipped.
Turn integers back into real numbers
Add dq <양자.npy> <출력.npy> <scale> <zero_point> <mode> (the placeholders are the quantized array and the output) so that it turns an integer array back into a float32 real-number array. The response must have min, max and count.
The definition is (q - zero_point) * scale. In this direction there is neither rounding nor saturation — a value that was cut off earlier does not come back to life here. Raise the integer array to float32 first and then subtract, so you avoid negatives wrapping around in uint8.
The round-trip error and its upper bound
Add roundtrip <입력.npy> <mode> [lo hi] (the placeholder is the input) so that it quantizes, immediately dequantizes, and outputs the maximum absolute error and the upper bound together. bound is half the scale and within_bound is whether the error is within it.
Quantization moves a value to the nearest point on a grid of the scale spacing. Since the grid spacing is the scale, at most it is half. Because of floating point it can overshoot by a tiny amount, so leave a little slack in the comparison. This upper bound holds only when the values are within the range.
What breaks when you narrow the range
Round-trip data/spread.npy once with the observed range as it is and once narrowed to -1.0 1.0, and write the two results in /root/onnxq-scale/saturate.json as full, narrow and error_ratio.
A narrow range makes the scale fine — that in itself is a good thing. The problem is that values that went outside the range stick to the wall. If the error greatly exceeds the upper bound, the cause is the range, not the rounding. error_ratio is the narrowed side's error divided by the original error.
If you apply symmetric to data that is only 0 or above
Apply asym-u8 and sym-i8 to data/positive.npy and write asym, sym and scale_ratio in /root/onnxq-scale/symmetry.json. Each item must have scale, zero_point, max_abs_error and usable_levels.
usable_levels is the number of grid steps that fall within the range [lo, hi] — you get it by dividing the range width by the scale, adding 1 and rounding. For data that is all 0 or above, asymmetric should give 256 and symmetric 128. Check whether the scale ratio becomes the error ratio as it is.
Confine one outlier
Add channel <입력.npy> <axis> (the placeholder is the input) so that it compares the per-channel scale and the whole-tensor scale with symmetric int8, and write the result of running data/weights.npy with axis 0 in /root/onnxq-scale/channel.json.
See one axis as the channels and find the min and max separately for each channel to make the scale. The per-channel scale is a 1-dimensional array, and for broadcasting you must reshape it so that only that axis has length. channels_improved is the number of channels whose per-channel error became smaller than the whole-tensor error. The key point is that the channel holding the outlier does not get better.
What you measured, on one page
Write bound_rule, spread, positive_asym_over_sym, channel_gain and saturation_ratio in /root/onnxq-scale/report.json, and write /root/onnxq-scale/report.md in the four sections ## 스케일과 영점은 어디서 나오나 ## 왕복 오차의 상한 ## 대칭과 비대칭 ## 이상값 하나가 하는 일 (the Korean headings mean "Where the scale and zero point come from", "The upper bound of the round-trip error", "Symmetric and asymmetric" and "What one outlier does").
For channel_gain, it is more honest to base it on the most improved channel rather than the whole-tensor error divided by the per-channel error. The channel holding the outlier stays the same, so the overall maximum error is almost the same. In the report, write along with each number one line on what that number refers to.