Is This Difference Reproducible? Measuring Accuracy Loss
Goal
Run the three — the FP32 original, the previous deployment and the integer candidate — on the same test samples, leave the predictions and output values, and build a tool acccmp.py that judges whether an accuracy difference is a reproducible difference. You measure the interval, the paired comparison, the flipped samples and the output difference separately, and reach a verdict with two baselines.
Why it matters
"Accuracy dropped by 0.3%" is not in itself a basis for judgment. On 400 samples, the 95% interval of a single accuracy is close to 8 percentage points wide. With the value from measuring two models separately and subtracting, you cannot speak of a difference smaller than this. What lets you speak is the same samples. If the two models saw the same inputs, the samples both got right and the samples both got wrong cannot be used for the judgment, and only the samples one got right are information. The number of those that split is the real denominator. The baseline is not one either. The FP32 original answers how much quantization shaved off, and the previous deployment answers whether it gets worse than what is running now. The two answers can point in different directions, so you write both. And output difference and the task metric are different. Even if the maximum absolute error comes out large, the accuracy can stay the same, and the reverse also happens. That is because samples far from the decision boundary do not change their answer even if the output swings. The grader does not trust the numbers you wrote. Each time it sets up its own prediction table with a different seed and number of samples, actually runs your tool, and redoes the same computation to check against it. Step 1 rebuilds the test inputs with the seed you left and directly runs your model.
Steps
- Create and run /root/acc/gen_eval.py to make the three models, /root/acc/preds.json and /root/acc/eval_seed.json.
- Create
summaryin /root/acc/acccmp.py so that it outputs the number of samples and the accuracy per run. - Add
intervalso that it outputs the 95% interval of a single accuracy as a Wilson score interval. - Add
pairedso that it outputs the counts split in 2x2 on the same samples and the paired comparison statistic. - Add
flipsso that it counts only the flipped predictions by direction. - Add
outputsso that it outputs the output value difference (maximum absolute error, mean absolute error and cosine). - Add
baselineso that it compares against each of the two baselines, and write the verdict for your data in /root/acc/verdict.json. - Write /root/acc/accuracy_report.md in four sections.
Notes
- Python is
/opt/onnx-lab/bin/python. The systempython3has neither numpy nor onnx. - Execution contract:
/opt/onnx-lab/bin/python /root/acc/acccmp.py <명령> <preds.json> [인자...](the placeholders are the command and the arguments). On success the exit code is 0, if the file is missing it is 3, and if the usage is wrong it is 2. The answer is output as one JSON blob on standard output. - The shape of
preds.json:{"labels": [정수...], "runs": {"fp32": {"pred": [정수...], "logits": [[실수...], ...]}, "prev": {...}, "int8": {...}}}(where the placeholders are integers and real numbers). Write logits as they are without rounding. - The shape of
eval_seed.json:{"seed": 정수, "n": 400, "features": 16, "classes": 10, "label_sigma": 1.0}(where the placeholder is an integer). - The test input is
numpy.random.default_rng(seed).standard_normal((n, features)).astype(numpy.float32). The answer table is the argmax of the FP32 output (raised to float64) plusnumpy.random.default_rng(seed + 1).standard_normal(모양) * label_sigma(where the placeholder is the shape). The grader rebuilds it with this rule and directly runs your model. - The model has the input name
x, the input shape[None, 16]and 10 output classes.prev.onnxis made withquantize_dynamicandint8.onnxwithquantize_static(quant_format=QDQ). The three file names are/root/acc/fp32.onnx,/root/acc/prev.onnxand/root/acc/int8.onnx. summaryresponse:{"n": 정수, "runs": {이름: {"correct": 정수, "accuracy": 실수}}}(where the placeholders are an integer, a name and a real number).intervalresponse:{"run", "n", "correct", "accuracy", "center", "low", "high", "half_width"}. It computes the Wilson score interval with z = 1.96. The center is not p but(p + z^2/(2n)) / (1 + z^2/n), and the half-width isz/(1 + z^2/n) * sqrt(p(1-p)/n + z^2/(4n^2)).paired <a> <b>response:{"a","b","n","both_correct","only_a","only_b","both_wrong","accuracy_a","accuracy_b","delta","discordant","statistic","significant"}. delta is b minus a. discordant is the sum of only_a and only_b. statistic is, by the yardstick this lab fixes,(|only_a - only_b| - 1)^2 / discordant, and is 0.0 if discordant is 0. significant is whether statistic is greater than 3.841459.flips <a> <b>response:{"a","b","changed","to_wrong","to_right","changed_both_wrong","top_flip"}. changed is the number of samples whose prediction differs, to_wrong is the number where a is right and b is wrong, to_right is the opposite, and changed_both_wrong is the number where both are wrong but the answer changed. top_flip is[a 예측, b 예측, 건수]of the most frequent flip (the placeholders are the a prediction, the b prediction and the count); in a tie, the one with the smaller a prediction and b prediction, and null if there is no flip.outputs <a> <b>response:{"a","b","max_abs_error","mean_abs_error","cos_min","cos_mean","argmax_changed"}. The mean absolute error is the mean over all elements, and the cosine is computed for each sample and gives the minimum and the mean. If the denominator is 0, that sample's cosine is taken as 1.0.baseline <run>response:{"run","baselines","vs","worst"}. baselines is the sorted names other than run, and each item of vs is{"delta","only_a","only_b","discordant","statistic","significant"}, the value with that baseline as a and run as b. worst is the{"baseline","delta"}of the baseline with the smallest delta.- The shape of
verdict.json:{"candidate","n","accuracy","delta_vs_fp32","delta_vs_prev","half_width","discordant_vs_fp32","statistic_vs_fp32","significant_vs_fp32","decision","reason"}. half_width is the candidate's Wilson half-width. decision follows the rule this lab fixes: if significant_vs_fp32 is true and delta_vs_fp32 is less than 0, it ishold, otherwiseship. accuracy_report.mdhas the four sections## 무엇을 견주었나## 같은 표본에서 본 차이## 이 차이가 재현되는가## 판정과 한계(the Korean headings mean "What was compared", "The difference seen on the same samples", "Is this difference reproducible" and "Verdict and limits").- Official documents: ONNX Runtime quantization · ONNX Concepts · QuantizeLinear
- Common mistakes: subtracting two separately measured accuracies and reporting that, saying there is no difference because the intervals overlap, substituting the maximum absolute error for accuracy, and having only FP32 as the baseline.
- This lab does not measure time or throughput. This Mac is an emulation, so latency wobbles by up to a factor of two.
Run the three versions on the same samples and leave the predictions
Create and run /root/acc/gen_eval.py to make fp32.onnx, prev.onnx and int8.onnx under /root/acc and /root/acc/preds.json and /root/acc/eval_seed.json. The test inputs and the answer table follow the seed rule in the Notes section exactly.
Build the models by chaining MatMul + Add + Relu with onnx.helper. Make prev.onnx with quantize_dynamic and int8.onnx by passing a CalibrationDataReader to quantize_static. The reason you build the answer table from the FP32 output but add noise is that if the accuracy is 1.0, no difference is visible. Write logits as they are with float() and do not round them.
Gather the sample count and accuracy in one place
Create summary <preds.json> in /root/acc/acccmp.py so that it outputs the number of samples and the number of correct answers and accuracy per run as JSON.
Do not start with the totals; leave for each sample a list of true/false for whether it was right. The paired comparison and flip counting in the later steps are all done on this list. Put the run names in sorted.
Attach an interval to a single accuracy
Add interval <preds.json> <run> so that it outputs that run's accuracy with a 95% Wilson score interval attached. It holds run, n, correct, accuracy, center, low, high and half_width.
Fix z at 1.96. The center of a Wilson interval is not the sample ratio p but a value pulled slightly toward 0.5. That is why low and high do not have p in the middle. The reason to use this interval is that the width does not become 0 even when the accuracy is 0 or 1.
Compare in pairs on the same samples
Add paired <preds.json> <a> <b> so that it outputs the 2x2 table (both_correct, only_a, only_b, both_wrong) and delta, discordant, statistic and significant.
The samples both got right and the samples both got wrong cannot separate the two models. The denominator of the judgment is not n but the number that split. Use the formula in the Notes section for the statistic as it is, and set it to 0.0 if the number that split is 0.
Count only the flipped predictions by direction
Add flips <preds.json> <a> <b> so that it outputs changed, to_wrong, to_right, changed_both_wrong and top_flip.
changed is the sum of the other three. If this identity does not hold, you counted something twice or missed it. The number of cases where one wrong answer changed into another wrong answer has no effect on accuracy, but to the user's eyes it looks like the answer changed.
Measure the output difference and the task metric separately
Add outputs <preds.json> <a> <b> so that it outputs max_abs_error, mean_abs_error, cos_min, cos_mean and argmax_changed.
The output having moved a lot does not mean the decision changes. For samples far from the decision boundary, the answer stays the same even if the output swings. If you put the two values in one response and look at them side by side, that relationship comes into view.
Judge with two baselines
Add baseline <preds.json> <run> so that it compares using each of the other runs as a baseline, and write the verdict for your data in /root/acc/verdict.json with the keys from the Notes section.
The answers for the original and the previous deployment can point in different directions. worst is the baseline with the smallest delta, and in a tie the one whose name comes first. The decision rule is fixed in the Notes section, so copy it as it is, and in reason write in one line why it came out that way.
Write it so that the reader can judge
Write /root/acc/accuracy_report.md in the four sections ## 무엇을 견주었나 ## 같은 표본에서 본 차이 ## 이 차이가 재현되는가 ## 판정과 한계 (the Korean headings mean "What was compared", "The difference seen on the same samples", "Is this difference reproducible" and "Verdict and limits"). The sample count, the number that split and the verdict must be in the text as numbers and words.
When you write a number, write where the number came from. In the limits section, write what this verdict does not say — this lab measured neither time nor memory, and saw things on only one test sample.