TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Is This Difference Reproducible? Measuring Accuracy Loss

Continue in TT Lab

Goal

Run the three — the FP32 original, the previous deployment and the integer candidate — on the same test samples, leave the predictions and output values, and build a tool acccmp.py that judges whether an accuracy difference is a reproducible difference. You measure the interval, the paired comparison, the flipped samples and the output difference separately, and reach a verdict with two baselines.

Why it matters

"Accuracy dropped by 0.3%" is not in itself a basis for judgment. On 400 samples, the 95% interval of a single accuracy is close to 8 percentage points wide. With the value from measuring two models separately and subtracting, you cannot speak of a difference smaller than this. What lets you speak is the same samples. If the two models saw the same inputs, the samples both got right and the samples both got wrong cannot be used for the judgment, and only the samples one got right are information. The number of those that split is the real denominator. The baseline is not one either. The FP32 original answers how much quantization shaved off, and the previous deployment answers whether it gets worse than what is running now. The two answers can point in different directions, so you write both. And output difference and the task metric are different. Even if the maximum absolute error comes out large, the accuracy can stay the same, and the reverse also happens. That is because samples far from the decision boundary do not change their answer even if the output swings. The grader does not trust the numbers you wrote. Each time it sets up its own prediction table with a different seed and number of samples, actually runs your tool, and redoes the same computation to check against it. Step 1 rebuilds the test inputs with the seed you left and directly runs your model.

Steps

  1. Create and run /root/acc/gen_eval.py to make the three models, /root/acc/preds.json and /root/acc/eval_seed.json.
  2. Create summary in /root/acc/acccmp.py so that it outputs the number of samples and the accuracy per run.
  3. Add interval so that it outputs the 95% interval of a single accuracy as a Wilson score interval.
  4. Add paired so that it outputs the counts split in 2x2 on the same samples and the paired comparison statistic.
  5. Add flips so that it counts only the flipped predictions by direction.
  6. Add outputs so that it outputs the output value difference (maximum absolute error, mean absolute error and cosine).
  7. Add baseline so that it compares against each of the two baselines, and write the verdict for your data in /root/acc/verdict.json.
  8. Write /root/acc/accuracy_report.md in four sections.

Notes

Run the three versions on the same samples and leave the predictions

Create and run /root/acc/gen_eval.py to make fp32.onnx, prev.onnx and int8.onnx under /root/acc and /root/acc/preds.json and /root/acc/eval_seed.json. The test inputs and the answer table follow the seed rule in the Notes section exactly.

Build the models by chaining MatMul + Add + Relu with onnx.helper. Make prev.onnx with quantize_dynamic and int8.onnx by passing a CalibrationDataReader to quantize_static. The reason you build the answer table from the FP32 output but add noise is that if the accuracy is 1.0, no difference is visible. Write logits as they are with float() and do not round them.

Gather the sample count and accuracy in one place

Create summary <preds.json> in /root/acc/acccmp.py so that it outputs the number of samples and the number of correct answers and accuracy per run as JSON.

Do not start with the totals; leave for each sample a list of true/false for whether it was right. The paired comparison and flip counting in the later steps are all done on this list. Put the run names in sorted.

Attach an interval to a single accuracy

Add interval <preds.json> <run> so that it outputs that run's accuracy with a 95% Wilson score interval attached. It holds run, n, correct, accuracy, center, low, high and half_width.

Fix z at 1.96. The center of a Wilson interval is not the sample ratio p but a value pulled slightly toward 0.5. That is why low and high do not have p in the middle. The reason to use this interval is that the width does not become 0 even when the accuracy is 0 or 1.

Compare in pairs on the same samples

Add paired <preds.json> <a> <b> so that it outputs the 2x2 table (both_correct, only_a, only_b, both_wrong) and delta, discordant, statistic and significant.

The samples both got right and the samples both got wrong cannot separate the two models. The denominator of the judgment is not n but the number that split. Use the formula in the Notes section for the statistic as it is, and set it to 0.0 if the number that split is 0.

Count only the flipped predictions by direction

Add flips <preds.json> <a> <b> so that it outputs changed, to_wrong, to_right, changed_both_wrong and top_flip.

changed is the sum of the other three. If this identity does not hold, you counted something twice or missed it. The number of cases where one wrong answer changed into another wrong answer has no effect on accuracy, but to the user's eyes it looks like the answer changed.

Measure the output difference and the task metric separately

Add outputs <preds.json> <a> <b> so that it outputs max_abs_error, mean_abs_error, cos_min, cos_mean and argmax_changed.

The output having moved a lot does not mean the decision changes. For samples far from the decision boundary, the answer stays the same even if the output swings. If you put the two values in one response and look at them side by side, that relationship comes into view.

Judge with two baselines

Add baseline <preds.json> <run> so that it compares using each of the other runs as a baseline, and write the verdict for your data in /root/acc/verdict.json with the keys from the Notes section.

The answers for the original and the previous deployment can point in different directions. worst is the baseline with the smallest delta, and in a tie the one whose name comes first. The decision rule is fixed in the Notes section, so copy it as it is, and in reason write in one line why it came out that way.

Write it so that the reader can judge

Write /root/acc/accuracy_report.md in the four sections ## 무엇을 견주었나 ## 같은 표본에서 본 차이 ## 이 차이가 재현되는가 ## 판정과 한계 (the Korean headings mean "What was compared", "The difference seen on the same samples", "Is this difference reproducible" and "Verdict and limits"). The sample count, the number that split and the verdict must be in the text as numbers and words.

When you write a number, write where the number came from. In the limits section, write what this verdict does not say — this lab measured neither time nor memory, and saw things on only one test sample.