Fact-check a smaller model’s performance claims
In one line
Measure the model file size and the execution performance separately, and state the scope of the evidence for deployment explicitly.
Why this was needed
If you look at a shrunken model and report "memory went down and it got faster everywhere", what is missing? A file is a number of bytes on disk, and a running process also includes Python, the inference engine and working buffers. A deployment decision needs not only accuracy but also what was measured in what environment. This lab is not a contest that guarantees quantization wins but practice in deciding whether to deploy with evidence.
How it works
The provided bench.py keeps model loading and input preparation outside the timing and measures the engine.run call. So the reported latency is not the full response time of a web request. The CPU inference threads are fixed at 1, and in 3 rounds of independent processes, for each batch size it measures 100 times after 10 warmups. The batch sizes are 1 and 32. The order of running the models also alternates in each round. The measured values are in microseconds per batch, and a total of 300 raw timings remain for each batch.
When implementing latency_stats(samples_us, batch_size), p50 is the median of the sorted samples, and p95 is the ceil(0.95×n)-th value by the nearest-rank method. Subtract 1 here for the Python index. The median of an even number of samples is the mean of the middle two. You do not average the per-round p95; you merge the raw samples and compute again. Throughput is batch_size×n×1,000,000/sum(samples_us). If you simply invert the batch-32 time into requests per second, the number of images is off by a factor of 32.
Memory is recorded by converting VmHWM in Linux /proc/self/status into bytes. This is the peak resident memory of the whole measuring process, and it is neither the model's memory alone nor GPU memory. In one measurement in this development container, even when the model file shrank from about 70KB to 24KB, the process HWM was similar at about 62MB. Time varies with shared CPU load, instruction set, runtime and so on, so those numbers are not pinned as the grading answer.
benchmark.json records the SHA-256 of the model and the data together. If you converted the model again, you must measure again. The hash linkage and the statistics checks find stale reports or calculation errors, but they are not a device for proving that the student did not write the times by hand. The raw samples are a reviewable experiment record, not a tamper-proof certificate.
Questions asked in a practical review meeting
Before asking "how much faster did it get?", ask "from where to where did you measure?". A cold start that measures up to reading the file and loading the model and the run call of an already-prepared session are different situations that the user experiences. A person who opens an app once a day and a device that keeps processing video in the same session have different metrics that matter. This lab measures prepared CPU inference, which is part of the latter, and does not include network, image preprocessing or result delivery latency.
If the p95 of three rounds are 10, 20 and 100 respectively, can you write their mean 43.3 as the overall p95? A percentile is determined by the number and order of the data points. With only three round summaries, you cannot reconstruct how many slow samples there were on which side. That is why the helper keeps all the raw timings and the student function re-sorts the merged 300. Conversely, if you throw away all the raw timings and store only summaries, the basis for re-examining outliers disappears too.
Also do a hand calculation with units attached. With batch size 32 and 300 repetitions, the total number of images is 9,600. If the time sum is 3,000,000 microseconds, that is 3 seconds, so the throughput is 3,200 images per second. If you use the number of batch calls as the numerator, it becomes 100 calls per second, and that value itself is not wrong. The mistake lies in naming it image throughput. The keys and units of the report must pin down the meaning of the formula.
A model that is fast on average but occasionally shows a long latency may not suit work with deadlines. However, this lab's p95 does not guarantee a hard real-time upper bound. There are load and device conditions not yet observed. Power, heat and long-duration stability also need separate tests. Leaving target_benchmark_required as true is not a phrase for leaving the work unfinished but a handover that precisely passes on the parts you cannot claim with this evidence.
Finally, write down what you must redo when you convert the model again. It is not just writing a new hash; you must rerun the inference regression and the measurement. If you attach the new file's hash to the old timing samples, the linkage format is right but the experiment record is false. Automatic grading cannot watch every behavior. Leaving the reproduction script, the raw results and the run conditions together so that a colleague can check again the same way is the basis of practice.
What you will do in the next lab
In step 7 you implement the statistics function and run the provided measurement helper. In release.json of step 8 you record the deployment conditions of this experiment: a model file of 32KiB or less and an accuracy loss of 1 percentage point or less. This does not mean the whole process runs within 32KiB. target_benchmark_required is true and universally_faster is false. Verification of latency, memory and power on a real device remains separate. Keep the sources and the report before the end of the lab. After the session ends, the working files are not kept.