TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

The AI Diet Gone Wrong

Continue in TT Lab

Goal

Convert the handwriting classifier into a real INT8 ONNX and reject the wrongly calibrated model. You measure accuracy, latency and memory on a Linux CPU, and do not do MCU, NPU, GPU or real power measurement.

Why it matters

Even if a model is small, if the input unit and calibration data are wrong, it runs as if normal and gives wrong answers. Only if you check file size, accuracy, execution kernels, latency and memory separately is there a basis for the deployment decision. You need the basics of Python, NumPy and JSON. The model and dependencies are provided, so no internet installation is needed.

Steps

  1. Write the input contract for the digit reader — save features=64, divisor=16, train=1077, calibration=360, test=360 as integers in contract.json. Check for yourself the shapes and ID counts of the provided digits.npz arrays. Do not put in other keys.
  2. Fix the input that used to be divided by 255 — implement normalize(pixels) in preprocess.py. It converts a uint8 N×64 array to float32 and returns a new array divided by 16. It keeps the input shape and does not modify the original array. You do not need to run the file directly.
  3. Keep evaluation samples from getting mixed into the calibration data — in the ids key of calibration.json, save the 360 integer IDs of the provided ids_calibration without duplicates. The order is free, but do not mix in training or evaluation IDs. Do not confuse row positions with original IDs.
  4. Put the AI on a diet — write and run quantize.py to create candidate.onnx. Copy the provided fp32.onnx as source.onnx, then convert the x_calibration corresponding to the IDs of calibration.json with normalize and have a CalibrationDataReader read it. The input name is pixels and the shape of one sample is 1×64. Use quantize_static with QDQ, activation QInt8, weight QInt8 and per_channel=True. It checks the integer weights of the three linear layers, real integer kernel execution, and an accuracy loss of 0.01 or less relative to FP32.
  5. Record the direction of the wrong digits — implement summarize(predictions, labels) in metrics.py. It takes two non-empty lists of integers 0–9 of the same length and returns only count, correct, accuracy and confusion. count and correct are integers, accuracy is a 0–1 ratio, and confusion is a 10×10 integer array of actual class rows × predicted class columns. The trace is correct.
  6. Reject the model that is small but has become dumb — run CPU inference of FP32, candidate.onnx and the provided bad-calibration.onnx on the same x_test/16, compare with y_test, and write regression.json. Save the step 5 metrics for each of fp32, int8 and bad, and for reject_bad write as a boolean whether the value of the FP32 accuracy minus the bad accuracy is greater than 0.01. The string true is not a boolean. Do not write the accuracy from memory; record the execution results.
  7. Do not say it got faster just because it got smaller — implement latency_stats(samples_us, batch_size) in benchmark.py. For a list of positive finite times, p50_us is the median, p95_us is the ceil(0.95×n)-th value by nearest-rank, and samples_per_second is batch_size×n×1e6/time sum. Run the provided bench.py --work /root/quantization to create benchmark.json. For each model and batch 1/32 it needs 300 raw timings, 3 rounds, threads=1, unit=us/batch, the SHA-256 of the data and model, and a positive peak_rss_bytes sourced from proc/VmHWM. The summary statistics must equal the raw samples.
  8. Deploy the evidence and the limits together — save model=candidate.onnx, sha256=the current candidate hash, max_file_bytes=32768, max_accuracy_drop=0.01, target_benchmark_required=true and universally_faster=false in release.json. Do not put in other keys. The actual model must be 32KiB or less, and the accuracy regression, the statistics and the link to the measured model are also checked again. If you changed the candidate, regenerate the step 6–7 results too.

Notes

All written and generated files are under /root/quantization. First run mkdir -p /root/quantization and then cd /root/quantization. Python uses /opt/onnx-lab/bin/python. Example of running a script: /opt/onnx-lab/bin/python quantize.py. The provided materials folder is /opt/lab/quantization. digits.npz holds the arrays x_train/x_calibration/x_test, y_train/y_calibration/y_test and ids_train/ids_calibration/ids_test. fp32.onnx is the baseline model, bad-calibration.onnx is the comparison group to reject, and reference.json, splits.json and DATA-LICENSE.json are the contract, the splits and the source. These binary materials are in the image, so they are not rewritten as text. The candidate ONNX is also generated by the converter. The step 7 command: /opt/onnx-lab/bin/python /opt/lab/quantization/bench.py --work /root/quantization. benchmark.json is generated by the measurement helper, so do not fabricate and insert samples by hand. Grading only checks the numeric calculations and the hash linkage and does not certify the truthfulness of the measurement. The submitted sources and JSON are each regular files of 64KiB or less, and the candidate model is 128KiB or less at check time and 32KiB or less at final deployment. The student functions are called in a separate process, so do not put in output logs. If you need a function from an earlier step, refer to the solution sheet and run script that support importing. The function check time limit is 10 seconds per call. If you run short of lab time, extend it with the +time button, and keep the files separately before it ends. After the session ends, the files are not kept.

Write the input contract for the digit reader

Save features=64, divisor=16, train=1077, calibration=360, test=360 as integers in contract.json. Check for yourself the shapes and ID counts of the provided digits.npz arrays. Do not put in other keys.

Look at the array names with np.load(..., allow_pickle=False) and data.files, and check the contract with shape and len.

Fix the input that used to be divided by 255

Implement normalize(pixels) in preprocess.py. It converts a uint8 N×64 array to float32 and returns a new array divided by 16. It keeps the input shape and does not modify the original array. You do not need to run the file directly.

Do the dtype conversion before the division. If you use 255 just because it is an image, it differs from this model's input.

Keep evaluation samples from getting mixed into the calibration data

In the ids key of calibration.json, save the 360 integer IDs of the provided ids_calibration without duplicates. The order is free, but do not mix in training or evaluation IDs. Do not confuse row positions with original IDs.

A NumPy integer array can be turned into a list that can be written to JSON with tolist().

Put the AI on a diet

Write and run quantize.py to create candidate.onnx. Copy the provided fp32.onnx as source.onnx, then convert the x_calibration corresponding to the IDs of calibration.json with normalize and have a CalibrationDataReader read it. The input name is pixels and the shape of one sample is 1×64. Use quantize_static with QDQ, activation QInt8, weight QInt8 and per_channel=True. It checks the integer weights of the three linear layers, real integer kernel execution, and an accuracy loss of 0.01 or less relative to FP32.

get_next() returns one input dictionary at a time and None when it is finished. You cannot write a temporary file next to the protected original, so use a working copy.

Record the direction of the wrong digits

Implement summarize(predictions, labels) in metrics.py. It takes two non-empty lists of integers 0–9 of the same length and returns only count, correct, accuracy and confusion. count and correct are integers, accuracy is a 0–1 ratio, and confusion is a 10×10 integer array of actual class rows × predicted class columns. The trace is correct.

If an actual 1 is predicted as 2, [1][2] goes up. Even if you flip the rows and columns the accuracy is the same, so check the off-diagonal cells.

Reject the model that is small but has become dumb

Run CPU inference of FP32, candidate.onnx and the provided bad-calibration.onnx on the same x_test/16, compare with y_test, and write regression.json. Save the step 5 metrics for each of fp32, int8 and bad, and for reject_bad write as a boolean whether the value of the FP32 accuracy minus the bad accuracy is greater than 0.01. The string true is not a boolean. Do not write the accuracy from memory; record the execution results.

Use ORT CPUExecutionProvider with intra/inter threads set to 1, and pass the argmax(axis=1) of the output to summarize.

Do not say it got faster just because it got smaller

Implement latency_stats(samples_us, batch_size) in benchmark.py. For a list of positive finite times, p50_us is the median, p95_us is the ceil(0.95×n)-th value by nearest-rank, and samples_per_second is batch_size×n×1e6/time sum. Run the provided bench.py --work /root/quantization to create benchmark.json. For each model and batch 1/32 it needs 300 raw timings, 3 rounds, threads=1, unit=us/batch, the SHA-256 of the data and model, and a positive peak_rss_bytes sourced from proc/VmHWM. The summary statistics must equal the raw samples.

The helper does the measuring and you write the statistics function yourself. The run command for this step is in the Notes section. Do not confuse the file size with the whole-process HWM.

Deploy the evidence and the limits together

Save model=candidate.onnx, sha256=the current candidate hash, max_file_bytes=32768, max_accuracy_drop=0.01, target_benchmark_required=true and universally_faster=false in release.json. Do not put in other keys. The actual model must be 32KiB or less, and the accuracy regression, the statistics and the link to the measured model are also checked again. If you changed the candidate, regenerate the step 6–7 results too.

Hash the model bytes with hashlib.sha256. This experiment's file budget is not an RSS budget, and you must not write that verification of target device performance was completed.