Count What You Actually Got: Execution Providers and Error Accumulation
Goal
Build a tool eprun.py that measures the execution provider you requested and the provider the session actually got, counts how many pieces the graph splits into depending on the list of supported operators, and makes a table of how quantization error grows at each layer boundary.
Why it matters
An execution provider is something you request, not something you are guaranteed. If you measure directly in this environment, even if you put a provider that does not exist in the list, no exception occurs and it falls to the CPU and runs as it is. So the misconception "we are using an accelerator" goes on for a long time without leaving any trace in the log. There is only one way to check — ask the session. A provider does not take charge of the whole model but takes only the nodes it can handle. So if the support list is cut along the graph order, the model splits into several pieces and values are handed over between pieces. If the pieces split differently, the order of computation changes, and even with the same input to the same model, different machines can produce different numbers. The error side needs the same attitude. Quantization error newly arises at each layer and the error of an earlier layer goes into the next layer. But it does not grow evenly per layer, so you must not pass it off as a story; you have to measure at each layer boundary and make a table to see which layer is the problem. This Pod has only two providers (AzureExecutionProvider and CPUExecutionProvider). You do not treat an accelerator that does not exist as if it did, and deal only with what can be observed here. The grader actually runs your tool each time with a different model, seed and support list, redoes the same computation and checks against it.
Steps
- Create and run /root/ep/gen_ep.py to make /root/ep/fp32.onnx and /root/ep/int8.onnx. There are at least five layers.
- Create
providersin /root/ep/eprun.py and write the results per request list in /root/ep/providers.json. - Add
partitionso that it counts how many pieces the graph would split into when a list of supported operators is given. - Add
exposeto build a version with the layer-boundary tensors taken out as graph outputs, and save /root/ep/fp32_exposed.onnx and /root/ep/int8_exposed.onnx. - Add
layersso that it outputs the maximum absolute error, relative error and minimum cosine at each layer boundary. - Add
growthso that it outputs the ratio at which the error grows as layers stack, and write the result for your model in /root/ep/growth.json. - Add
settingsso that it outputs the results of running the same model and the same input with only the settings changed, and write it in /root/ep/settings.json. - Write /root/ep/ep_report.md in four sections.
Notes
- Python is
/opt/onnx-lab/bin/python. The systempython3has no onnxruntime. - Execution contract:
/opt/onnx-lab/bin/python /root/ep/eprun.py <명령> <인자...>(the placeholders are the command and the arguments). On success the exit code is 0, if the file is missing it is 3, and if the usage is wrong it is 2. The answer is output as one JSON blob on standard output. The EP list and the list of supported operators are joined with commas and given as one argument. - The model has the input name
xand the input shape[None, 특징수](the placeholder is the number of features), and stacks MatMul, Add and Relu at least five layers deep.int8.onnxis made withquantize_static(quant_format=QDQ). providers <model> <목록> [...]response (the placeholder is the list):{"model", "available", "cases"}. available isonnxruntime.get_available_providers()as it is. Each item of cases is{"requested", "actual", "missing", "fell_back", "error"}, where actual is the session'sget_providers(), missing is the names requested but not in available, fell_back is whether actual differs from requested, and error is the exception name if the session could not be created (otherwise null).- If you request a provider that does not exist, onnxruntime prints a fallback notice to standard output. If you leave it, it gets mixed into the answer JSON, so redirect standard output to standard error while creating the session, or at least make the JSON come on the last line.
providers.jsonis the response above with onenoteline added. Include both a request naming only providers that exist and a request mixing in a provider that does not exist.partition <model> <지원연산자목록>response (the placeholders are the model and the list of supported operators):{"model","supported","nodes","claimed","partition_count","handoffs","partitions"}. It sweeps in the order the nodes are written in the graph, marks an operator in the support list asepand otherwisecpu, and groups a stretch where the same owner continues as one piece. Each item of partitions is{"owner","nodes","size"}, and handoffs is the number of pieces minus 1 (0 if there are no pieces). This is a simulation, not the actual placement.expose <model> <out.onnx>response:{"model","out","boundaries","outputs"}. The layer boundaries are the first input tensors that MatMul and Gemm nodes take in, gathered in graph order, with the graph's first output attached at the end. A name that is already an output is not added again. If you try to take out an integer tensor as float in an integer version, the session will not open, so keep to the rule of taking the take-in-side tensor as the boundary.layers <fp32> <int8> <seed>response:{"seed","rows","layer_count","layers"}. The input isnumpy.random.default_rng(seed).standard_normal((64, 특징수)).astype(numpy.float32)(where the placeholder is the number of features) and rows is 64. Create the session withproviders=["CPUExecutionProvider"]and leave the other settings at their defaults. Each item of layers is{"index","fp32","int8","max_abs_error","scale","rel_error","cos_min"}, where scale is the maximum absolute value of the FP32-side tensor, rel_error is the maximum absolute error divided by scale (0.0 if scale is 0), and cos_min is the minimum of the cosines computed per row (1.0 for a row whose denominator is 0). The boundaries of the two graphs are paired by position order.growth <fp32> <int8> <seed>response:{"seed","layer_count","rel_error","ratios","monotone","total_growth","max_ratio"}. The k-th of ratios is the relative error of layer k divided by that of the layer before it, and null if the earlier layer is 0. monotone is whether the relative error never decreased, total_growth is the last divided by the first (null if the first is 0), and max_ratio is the{"index","ratio"}of the largest ratio, and in a tie the earlier position.settings <model> <seed>response:{"model","seed","baseline","cases","all_equal"}. The baseline isORT_DISABLE_ALLwith intra/inter threads 1, and cases holds three cases in this order: (ORT_ENABLE_ALL, 1), (ORT_DISABLE_ALL, 2) and (ORT_ENABLE_ALL, 2). Each item is{"level","intra_op","max_abs_diff","bitwise_equal"}, and baseline is{"level","intra_op","max_abs_output","mean_output"}.growth.jsonandsettings.jsonare the responses above for your model as they are. You choose the seed and that value goes into the file.ep_report.mdhas the four sections## 무엇을 요청했고 무엇을 쥐었나## 그래프가 쪼개지는 자리## 층을 거듭하며 자란 오차## 이 기계에서만 참인 것(the Korean headings mean "What was requested and what was obtained", "Where the graph splits", "The error that grew as layers stacked" and "What is true only on this machine").- Official documents: Execution Providers · Python API · Graph optimizations · Thread management · ONNX Concepts
- Common mistakes: leaving the request list as it is in the log, believing that putting in an EP that does not exist will raise an exception, trying to take out an intermediate tensor of an integer version as float so that the session does not open, and assuming the error grows monotonically per layer.
- This lab does not measure time or throughput. This Mac is an emulation, so latency wobbles by up to a factor of two.
Build a model with several layers stacked and its integer version
Create and run /root/ep/gen_ep.py to make /root/ep/fp32.onnx and /root/ep/int8.onnx. There must be at least five linear operators.
With two or three layers, the picture of the error growing does not show. Stack at least five layers. Static quantization needs a CalibrationDataReader, and the calibration data should be drawn from the same distribution as the test input but does not need to be the same samples.
Set what you requested beside what you got
Create providers <model> <목록> [...] (the placeholder is the list) in /root/ep/eprun.py, and write in /root/ep/providers.json the results of a request naming only providers that exist and a request mixing in a provider that does not exist.
You might think an exception will be raised if you put in a provider that does not exist, but in this environment it is not so. Put it in yourself and measure what happens. session.get_providers() is the answer, and the request list is not the answer. There might be a version that raises an exception, so also keep the exception name. And since the fallback notice comes out on standard output, block it so that the answer does not get mixed — contextlib.redirect_stdout is useful.
Count how many pieces the graph splits into
Add partition <model> <지원연산자목록> (the placeholders are the model and the list of supported operators) so that it simulates how the graph would split if there were a provider that supports only those operators. It outputs the number of pieces and the number of handoffs together.
It is not imitating the actual placement but counting the spots where the support list is cut along the graph order. A stretch where the same owner continues is one piece. If you assume only MatMul is supported on the integer version, you can see right away how much the pieces increase because of Q/DQ.
Take out the layer-boundary tensors
Add expose <model> <out.onnx> so that it saves a version with the layer-boundary tensors appended as graph outputs, and make /root/ep/fp32_exposed.onnx and /root/ep/int8_exposed.onnx.
In the integer version, the output of QuantizeLinear is int8. If you declare it to be taken out as float, the session will not open. The tensor the linear operator takes in is float in both graphs, so it is safe. Actually open and run the saved version before moving on.
Measure the error at each layer boundary
Add layers <fp32> <int8> <seed> so that it feeds the same input made with the seed into the two versions and outputs the maximum absolute error, magnitude, relative error and minimum cosine at each layer boundary.
If you look only at the absolute error, the later layers always look worse, because the values themselves get larger. That is why you also output the relative error divided by the maximum absolute value of that layer's tensor. The boundary names of the two graphs differ, so pair them by position order.
Make a table of the growth ratios
Add growth <fp32> <int8> <seed> so that it outputs the relative error per layer and its ratios, whether it is monotone, the overall growth multiple and the layer with the biggest jump, and write the result for your model in /root/ep/growth.json.
If you measure, you also get stretches where the ratio is less than 1. That is because the activation function cuts off a part and the scale is set anew at each layer. That it is not monotone is itself the result, so write it as it is. If the earlier layer's relative error is 0, you cannot divide, so leave it null.
Run the same input changing only the settings
Add settings <model> <seed> so that it runs the same model and the same input changing only the optimization level and the number of threads and outputs the difference from the baseline. Write the result for your model in /root/ep/settings.json.
In this Pod, all four cases gave the same values. Write that result as it is — do not fabricate that they must differ. What matters is not the fact that they were the same but that you have a procedure for measuring whether they are the same or different. On other machines, a different answer can come out.
Write what is true only on this machine
Write /root/ep/ep_report.md in the four sections ## 무엇을 요청했고 무엇을 쥐었나 ## 그래프가 쪼개지는 자리 ## 층을 거듭하며 자란 오차 ## 이 기계에서만 참인 것 (the Korean headings mean "What was requested and what was obtained", "Where the graph splits", "The error that grew as layers stacked" and "What is true only on this machine"). The number of usable providers and the overall error growth multiple must be in it as numbers.
The last section is the core of this report. Write which of the values measured here are tied to this Pod's provider configuration and these settings. Also write that you did not measure time.