When Does a 0.3% Accuracy Drop Actually Mean Something?
In one line
An accuracy difference does not come from subtracting two numbers but from the number of samples and the number of samples that split. If you do not compare on the same samples, 0.3 percentage points means nothing.
Why this becomes a problem
The sentence that comes up most in the place where you decide whether to put a quantized model up is "accuracy dropped by 0.3%". You cannot decide anything from this sentence alone. If there are 300 samples, 0.3 percentage points is a difference of one sample. It is a size at which the direction could change if you drew the same data again.
Accuracy is a ratio, and if you measure a ratio on a sample, it wobbles as much as the sample wobbles. If the accuracy on 400 test samples is 0.78, the 95% interval of that single number is roughly 0.74 to 0.82. The width is 8 percentage points. A person who has seen this width does not take 0.3 percentage points as evidence.
But here another common mistake comes up. Saying there is no difference because the intervals overlap. That is also wrong. If you measured the two models on the same samples, you have much more usable information.
How to reason about it
The key is that the two models saw the same inputs. If, of 400, 380 are answered correctly by both models and 3 are wrong for both, those 383 are of no help in separating the two models. Only the remaining 17, the samples where one is right and the other is wrong, are information.
So you make a 2x2 table.
B 맞음 B 틀림
A 맞음 307 6 <- A 만 맞은 6건
A 틀림 8 79 <- B 만 맞은 8건
In this table the accuracy difference is (8 - 6) / 400 = 0.005. But what matters is the fact that the denominator is not 400 but the 14 that split. If 14 cases split 6 to 8, it is no different from flipping a coin 14 times and getting 6 to 8. The paired comparison statistic turns this judgment into a single number, and this lab uses a yardstick divided by the number of cases that split. If the number of split cases is small, the statistic does not get large however large the difference looks.
You may ask, why not just increase the samples? That is true, but it is expensive. The width of the interval shrinks in inverse proportion to the square root of the number of samples, so to halve the width you must quadruple the samples. To narrow the half-width of the 95% interval from 0.04 to 0.005 for a model with accuracy 0.78, you need about 26 thousand test samples. Attaching that many labels is not realistic, so before increasing the samples, use the paired comparison on the same samples first. A paired comparison wipes out the wobble the two models experience together, so with the same number of samples it catches a much smaller difference.
The baseline is not one either. Usually there are two. The FP32 original answers "how much did quantization shave off", and the previous deployment answers "does it get worse than what is running now". The answers to the two questions can point in different directions. A candidate that is worse than the original but better than the previous deployment actually occurs.
Two things you must separate in the field
Output difference and the task metric are different things. Maximum absolute error and cosine similarity measure how much the output vector moved, and accuracy measures whether the decision changed. The two do not move together. There are cases where the maximum absolute error is as large as 1.7 and yet the accuracy actually goes up. Samples far from the decision boundary do not change their answer even if the output swings greatly.
So the report writes down the three together. How much the output moved (maximum absolute error and cosine), how many decisions flipped (the number of flipped samples and direction), and as a result what happened to the task metric (accuracy and its interval). If you write only one, the reader imagines the rest.
There is one more thing to be careful about. If you use the same test samples repeatedly for several candidates, a candidate that happens to fit those samples well gets picked. If you measure ten candidates and choose the best one, the interval of that one is optimistic compared with when you first measured it. So after comparing several candidates, you need a procedure to measure the chosen candidate again on samples that have never been used. This lab deals with one candidate, but you must know that this problem follows the moment the candidates increase.
One more thing. When counting flipped samples, you must split by direction. The number of cases that were right and became wrong, the number that were wrong and became right, and the number where one wrong answer changed into another wrong answer. The third item has no effect on accuracy but is evidence that the model wobbled, and to the user's eyes it looks like "the answer changed".
What really matters in practice
- Write accuracy together with an interval. If you write it as a single number, the reader misjudges the precision.
- Compare two models on the same samples. Do not subtract two separately measured numbers.
- Write the number of samples that split too. That is the real denominator.
- Keep two baselines. The answers for the original and the previous deployment can differ.
- Write the output difference and the task metric separately. Do not substitute one for the other.
What you will do in the next lab
You build yourself the test data and the predictions of three versions (the FP32 original, the previous deployment and the integer candidate), and grow a tool step by step that measures accuracy intervals, a paired comparison table, flipped samples and output differences. The grader does not trust the numbers you wrote. Each time it sets up its own prediction table with a different seed, actually runs your tool, and redoes the same computation to check against it. In step 1, it rebuilds the test inputs with the seed you left and directly runs your model.