TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Same numbers, different units

Continue in TT Lab

In one line

Only when the input unit and the role of each sample are fixed does comparing models become an experiment.

Why this was needed

Imagine moving the digit reader of a parcel-sorting robot onto a small device. The model that read well on the development PC gets the digits wrong after deployment. The model file opens normally and there is no error log. A failure like this cannot be found just by checking that the program runs. Only if you fix the input unit and the data split first can you find the cause when the results later change.

In this course you receive an already-trained small handwriting classifier. The task is not a contest of training a new model but model conversion and verification. You need the basics of Python functions, lists, JSON and NumPy arrays. It runs on a Linux CPU without a GPU or a physical robot, and the speed or power consumption of a real MCU or NPU cannot be proven with this lab.

How it works

The provided digits data is an array of 8×8 images flattened into 64 features. Pixel values are 0–16. If you divide by 255 just because it is an image you commonly see, the input range differs from what it was in training. normalize(pixels) must convert to a float32 array, divide by 16 and keep the input shape. There is no need to modify the original array itself.

An image's pixel values and a sample ID are different things. An ID is an integer that identifies a row of the original data. In digits.npz, x_train, x_calibration and x_test hold the images, y_test holds the evaluation answers, and ids_train, ids_calibration and ids_test hold the IDs of each split. The provided split is 1,077 for training, 360 for calibration and 360 for evaluation. You do not redraw the split and use the provided IDs as they are.

Calibration data is the material that decides the converter's numeric range, and evaluation data is the material for judging the quality after conversion. If you pick calibration samples while looking at the evaluation answers and raise the score, the roles of the two get mixed. A case where even one evaluation ID gets into the calibration file is also leakage. Do not just count the IDs; check the set and the duplicates. However, a score from repeatedly using this small public evaluation set is not independent generalization performance in the field.

What it looks like in the field

A preprocessing function is a small API between the data team and the application team. Even if one side expects 0–1 real numbers and the other puts in 0–16 integers, the call can succeed as long as the array size is the same. So the contract needs not only the number of features but also the data type and the range of values. In real systems that handle camera and sensor input, the channel order, the time unit and the handling of missing values are added to this. This input deliberately leaves out that complexity. First you check how an error spreads silently when you get a single range wrong.

Try a small calculation on paper first. If three pixels are 0, 8 and 16, the correct normalization result is 0.0, 0.5 and 1.0. If you divide by 255, even the last value is only about 0.0627. From the program's point of view both are finite real numbers, so no exception occurs. The model has received an unusually dark input. This is why you must separate the fact that there is no error from the fact that the input is correct.

You must also distinguish a function that modifies the whole array from a function that returns a new array. If you change the original x_test, when you evaluate the next model in the same process, you may divide already-normalized values once more. The evaluation input of each model changes, so you end up wrongly explaining a performance difference as the fault of quantization. This normalize contract preserves the input array. The separate function test also checks the two end values by putting in, besides the actual samples, a row of all 0 and a row of all 16.

It helps to split ID selection into three questions: "count, duplicates, set". Answer each: did you include 360, did you not put in the same ID twice, and is it the same as the provided calibration set? The last question must pass even if the order is different. If you fix the order, you reject a normal implementation that reads the same data in a different order. If you simply fill in 0 through 359, you lose track of which row of the original an ID means. When reading calibration samples, it is enough to build a mapping table that finds the row position inside the calibration array from the original ID.

The already-trained FP32 model and the data are fixed, so you start from the same bytes. Even so, this data does not represent every style of handwriting. The reproducibility of public data and the representativeness of real customer input are different problems. Before deployment in the field you must collect and evaluate data with separate handwriting, shooting conditions and contamination, and you cannot skip that work on the strength of the high score in this course. After confirming the input and the split in the quiz right after this, you move on to the integrated lab of the last module.

What you will do in the next lab

In steps 1–3 you record the input contract in contract.json, implement the normalization function in preprocess.py, and save only the IDs for calibration in calibration.json. Try getting the unit wrong first and see which check catches it. The first step begins with a small success of reading the shape and length of the provided arrays.

Data sources: scikit-learn load_digits, UCI original data. The original data's authors and the CC BY 4.0 notice are also in the provided DATA-LICENSE.json.