TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

The model got smaller—and stopped reading digits

Continue in TT Lab

In one line

You confirm conversion success, integer kernel execution and classification quality with different checks.

Why this was needed

You reduced the file size of the digit reader, but the accuracy dropped greatly. The conversion command succeeded and the file extension is onnx. On what basis do you reject the deployment? The fact that the file is small, the fact that there are integer weights, the fact that real integer kernels ran, and the fact that it reads digits properly are each different evidence.

How it works

Quantization is a conversion that approximates real numbers with a narrow integer representation. It maps between integers and real numbers using a scale and a zero point, so information loss can occur. Static quantization decides the activation ranges by running calibration inputs in advance. A QDQ graph contains QuantizeLinear and DequantizeLinear. In this process, the activations and weights are converted to signed INT8, and per-channel ranges are used for the weights. For the detailed API, refer to the ONNX Runtime official guide.

The FP32 model in our fixture is a small MLP with three linear layers. You build a CalibrationDataReader that reads the 360 calibration images and passes each sample as a 1×64 float32 input named pixels. The converter may create a temporary file for shape inference next to the input ONNX. So you do not convert the protected /opt/lab/quantization/fp32.onnx directly; you copy it to your own working folder as source.onnx and use that. There is no need to make the original writable.

The provided bad-calibration.onnx is a failing comparison group made with a calibration input in which every pixel is 0. Run it on the same evaluation data as the normal calibrated candidate and look at the difference. Which digit was mistaken for which other digit shows up in the confusion matrix. In this text, rows are the actual class and columns are the predicted class. If an actual 1 was read as 2, you increase confusion[1][2]. Even if you flip the axes, the trace and the overall accuracy are the same, so if you verify only accuracy you miss the error.

For a model with QDQ applied also to the final output, the sum of the probabilities may not be exactly 1 because of rounding. The minimum sum observed in the provided normal conversion was about 0.992. The checker allows the 1/255 spacing of this conversion and the rounding error of 10 classes, but rejects negative values, non-finite values and sums that are far off. Saying there is a tolerance does not mean letting any output pass.

What it looks like in the field

Suppose you take part in a deployment review of a classifier. If a colleague says "I made a 24KB file", there is no need to reject it right away. Instead you can ask on what input and with what performance it worked. File size is evidence of storage space. The weight data type is evidence of the representation. A profile is evidence of the execution path. A metric obtained on evaluation data is evidence of quality within a limited scope. The core of this assignment is not to merge these four questions into one item.

A confusion matrix makes the accuracy number explainable. If the actual answers are [1, 2, 1, 0, 9] and the predictions are [2, 2, 1, 0, 8], three samples are correct. The accuracy is 3/5, and the two wrong samples are 1→2 and 9→8. If the code recorded them as 2→1 and 8→9, the ratio is right but the explanation in the report is backwards. This is why you separate the summarize function from the whole model and test it with short lists.

When checking an accuracy difference, you must also fix the comparison baseline. If the baseline FP32 itself receives a wrong input, both it and the candidate can give low accuracy while only the difference gets small. If you set only the criterion "the loss is small", a model that neither can read passes. The checker first confirms that the provided FP32 is 0.95 or higher on the evaluation set and then looks at the difference from the candidate. The preprocessing-only test, the baseline model check and the candidate comparison each take charge of a different kind of error.

Think about why you actually run the bad calibration model too. A program that always writes reject_bad in the report as true only imitates the appearance of rejecting. Only by recomputing and cross-checking the count, correct, accuracy and confusion matrix of the three models can you know on what basis the bad candidate was rejected. The comparison group's model size may be similar to the normal candidate's, so you must not use the byte count as a substitute for the quality metric. The file being small is not the culprit in this case but a separate observation.

If the CPU type changes, even for an INT8 model with identical bytes, the predictions of samples near a boundary can change. You do not memorize one machine's answer list during development and apply it to all run environments. You evaluate the baseline and the candidate together in the same run environment and check the difference. However, that does not mean unconditionally allowing that a difference arose. If it exceeds the loss limit you set, it is a failure. After reading the error direction in the quiz right after this, you build the report yourself in the integrated lab of the last module.

What you will do in the next lab

In steps 4–6 you build candidate.onnx yourself and run inference on the three models. summarize(predictions, labels) in metrics.py returns count, correct, accuracy and the 10×10 confusion. Accuracy is a 0–1 ratio. The loss limit of 1 percentage point relative to FP32 is a ratio of 0.01, and is different from meaning a relative 1% decrease. If FP32 is 0.96, a candidate of 0.95 is the boundary. You leave the actual results and the rejection judgment of the bad model in regression.json.