TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Build the Ruler Yourself: Scale, Zero Point, Round-Trip Error

Continue in TT Lab

Goal

Build qmath.py to compute the scale and zero point exactly as the definitions of QuantizeLinear and DequantizeLinear, implement quantization and dequantization, confirm that the round-trip error is within half the scale, and compare symmetric and asymmetric, and per-channel scale and whole-tensor scale, on the same data.

Why it matters

To pick where to fix when the accuracy of the int8 model a conversion tool produced drops, you have to know the arithmetic the tool does inside. The scale is the observed range divided by the number of grid steps, and the zero point is the integer position where the real number 0.0 lands. If 0 is not in the range, the 0 that padding and ReLU made comes back as a non-zero value. The round-trip error has an upper bound. As long as the value is within the range, the error does not exceed half the scale — because quantization moves a value to a grid of the scale spacing. So if the error greatly exceeds that half, it is the range, not the rounding, that is the problem. This one yardstick makes the cause judgment fast. The shape of the data also decides the method. The activations after a ReLU are all 0 or above, so if you apply symmetric int8, the grid on the negative side sits idle as a whole and the scale becomes twice as coarse. If there is one outlier in a weight matrix, the whole-tensor scale gets dragged to that value and everything else is squashed — if you split the channels, the damage is confined within that channel. The grader does not trust the numbers you wrote out. Each time it builds arrays with a different seed, actually runs your qmath.py, and checks against the result of putting the same input straight through ONNX's QuantizeLinear and DequantizeLinear operators.

Steps

  1. Create and run /root/onnxq-scale/gen_data.py to make three sets of arrays under /root/onnxq-scale/data, and implement params in /root/onnxq-scale/qmath.py.
  2. Add q to bring a real-number array down to an integer array.
  3. Add dq to turn an integer array back into real numbers.
  4. Add roundtrip so that it outputs the round-trip error and its upper bound together.
  5. Deliberately narrow the range to create saturation and write the result in /root/onnxq-scale/saturate.json.
  6. Apply symmetric and asymmetric to data that is all 0 or above and write it in /root/onnxq-scale/symmetry.json.
  7. Add channel to compare the per-channel scale and the whole-tensor scale and write it in /root/onnxq-scale/channel.json.
  8. Summarize on one page with /root/onnxq-scale/report.json and /root/onnxq-scale/report.md.

Notes

Turn the observed range into a scale and zero point

Create and run /root/onnxq-scale/gen_data.py to make three arrays under /root/onnxq-scale/data, and implement params <lo> <hi> <mode> in /root/onnxq-scale/qmath.py. If 0 is not in the range, you must put it in.

For asymmetric, divide (hi - lo) by 255, and the zero point is the integer position where the real number 0.0 lands — round qmin minus lo/scale. For symmetric, divide the larger of the left and right widths by 127 and the zero point is 0. Handle the cases where lo is greater than 0 or hi is less than 0 first.

Bring real numbers down to an integer grid

Add q <입력.npy> <출력.npy> <mode> [lo hi] (the placeholders are the input and the output) so that it brings a real-number array down to an integer array and saves the result as .npy. The response must have scale, zero_point, clipped, dtype and count.

The definition is saturate(round(x / scale) + zero_point). Divide, round, add the zero point, and cut to the two ends of the data type. If you swap the order, the answer changes — you must not add the zero point first and then round. Count the number of elements that were cut and give it as clipped.

Turn integers back into real numbers

Add dq <양자.npy> <출력.npy> <scale> <zero_point> <mode> (the placeholders are the quantized array and the output) so that it turns an integer array back into a float32 real-number array. The response must have min, max and count.

The definition is (q - zero_point) * scale. In this direction there is neither rounding nor saturation — a value that was cut off earlier does not come back to life here. Raise the integer array to float32 first and then subtract, so you avoid negatives wrapping around in uint8.

The round-trip error and its upper bound

Add roundtrip <입력.npy> <mode> [lo hi] (the placeholder is the input) so that it quantizes, immediately dequantizes, and outputs the maximum absolute error and the upper bound together. bound is half the scale and within_bound is whether the error is within it.

Quantization moves a value to the nearest point on a grid of the scale spacing. Since the grid spacing is the scale, at most it is half. Because of floating point it can overshoot by a tiny amount, so leave a little slack in the comparison. This upper bound holds only when the values are within the range.

What breaks when you narrow the range

Round-trip data/spread.npy once with the observed range as it is and once narrowed to -1.0 1.0, and write the two results in /root/onnxq-scale/saturate.json as full, narrow and error_ratio.

A narrow range makes the scale fine — that in itself is a good thing. The problem is that values that went outside the range stick to the wall. If the error greatly exceeds the upper bound, the cause is the range, not the rounding. error_ratio is the narrowed side's error divided by the original error.

If you apply symmetric to data that is only 0 or above

Apply asym-u8 and sym-i8 to data/positive.npy and write asym, sym and scale_ratio in /root/onnxq-scale/symmetry.json. Each item must have scale, zero_point, max_abs_error and usable_levels.

usable_levels is the number of grid steps that fall within the range [lo, hi] — you get it by dividing the range width by the scale, adding 1 and rounding. For data that is all 0 or above, asymmetric should give 256 and symmetric 128. Check whether the scale ratio becomes the error ratio as it is.

Confine one outlier

Add channel <입력.npy> <axis> (the placeholder is the input) so that it compares the per-channel scale and the whole-tensor scale with symmetric int8, and write the result of running data/weights.npy with axis 0 in /root/onnxq-scale/channel.json.

See one axis as the channels and find the min and max separately for each channel to make the scale. The per-channel scale is a 1-dimensional array, and for broadcasting you must reshape it so that only that axis has length. channels_improved is the number of channels whose per-channel error became smaller than the whole-tensor error. The key point is that the channel holding the outlier does not get better.

What you measured, on one page

Write bound_rule, spread, positive_asym_over_sym, channel_gain and saturation_ratio in /root/onnxq-scale/report.json, and write /root/onnxq-scale/report.md in the four sections ## 스케일과 영점은 어디서 나오나 ## 왕복 오차의 상한 ## 대칭과 비대칭 ## 이상값 하나가 하는 일 (the Korean headings mean "Where the scale and zero point come from", "The upper bound of the round-trip error", "Symmetric and asymmetric" and "What one outlier does").

For channel_gain, it is more honest to base it on the most improved channel rather than the whole-tensor error divided by the per-channel error. The channel holding the outlier stays the same, so the overall maximum error is almost the same. In the report, write along with each number one line on what that number refers to.