TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

Count What You Actually Got: Execution Providers and Error Accumulation

Continue in TT Lab

Goal

Build a tool eprun.py that measures the execution provider you requested and the provider the session actually got, counts how many pieces the graph splits into depending on the list of supported operators, and makes a table of how quantization error grows at each layer boundary.

Why it matters

An execution provider is something you request, not something you are guaranteed. If you measure directly in this environment, even if you put a provider that does not exist in the list, no exception occurs and it falls to the CPU and runs as it is. So the misconception "we are using an accelerator" goes on for a long time without leaving any trace in the log. There is only one way to check — ask the session. A provider does not take charge of the whole model but takes only the nodes it can handle. So if the support list is cut along the graph order, the model splits into several pieces and values are handed over between pieces. If the pieces split differently, the order of computation changes, and even with the same input to the same model, different machines can produce different numbers. The error side needs the same attitude. Quantization error newly arises at each layer and the error of an earlier layer goes into the next layer. But it does not grow evenly per layer, so you must not pass it off as a story; you have to measure at each layer boundary and make a table to see which layer is the problem. This Pod has only two providers (AzureExecutionProvider and CPUExecutionProvider). You do not treat an accelerator that does not exist as if it did, and deal only with what can be observed here. The grader actually runs your tool each time with a different model, seed and support list, redoes the same computation and checks against it.

Steps

  1. Create and run /root/ep/gen_ep.py to make /root/ep/fp32.onnx and /root/ep/int8.onnx. There are at least five layers.
  2. Create providers in /root/ep/eprun.py and write the results per request list in /root/ep/providers.json.
  3. Add partition so that it counts how many pieces the graph would split into when a list of supported operators is given.
  4. Add expose to build a version with the layer-boundary tensors taken out as graph outputs, and save /root/ep/fp32_exposed.onnx and /root/ep/int8_exposed.onnx.
  5. Add layers so that it outputs the maximum absolute error, relative error and minimum cosine at each layer boundary.
  6. Add growth so that it outputs the ratio at which the error grows as layers stack, and write the result for your model in /root/ep/growth.json.
  7. Add settings so that it outputs the results of running the same model and the same input with only the settings changed, and write it in /root/ep/settings.json.
  8. Write /root/ep/ep_report.md in four sections.

Notes

Build a model with several layers stacked and its integer version

Create and run /root/ep/gen_ep.py to make /root/ep/fp32.onnx and /root/ep/int8.onnx. There must be at least five linear operators.

With two or three layers, the picture of the error growing does not show. Stack at least five layers. Static quantization needs a CalibrationDataReader, and the calibration data should be drawn from the same distribution as the test input but does not need to be the same samples.

Set what you requested beside what you got

Create providers <model> <목록> [...] (the placeholder is the list) in /root/ep/eprun.py, and write in /root/ep/providers.json the results of a request naming only providers that exist and a request mixing in a provider that does not exist.

You might think an exception will be raised if you put in a provider that does not exist, but in this environment it is not so. Put it in yourself and measure what happens. session.get_providers() is the answer, and the request list is not the answer. There might be a version that raises an exception, so also keep the exception name. And since the fallback notice comes out on standard output, block it so that the answer does not get mixed — contextlib.redirect_stdout is useful.

Count how many pieces the graph splits into

Add partition <model> <지원연산자목록> (the placeholders are the model and the list of supported operators) so that it simulates how the graph would split if there were a provider that supports only those operators. It outputs the number of pieces and the number of handoffs together.

It is not imitating the actual placement but counting the spots where the support list is cut along the graph order. A stretch where the same owner continues is one piece. If you assume only MatMul is supported on the integer version, you can see right away how much the pieces increase because of Q/DQ.

Take out the layer-boundary tensors

Add expose <model> <out.onnx> so that it saves a version with the layer-boundary tensors appended as graph outputs, and make /root/ep/fp32_exposed.onnx and /root/ep/int8_exposed.onnx.

In the integer version, the output of QuantizeLinear is int8. If you declare it to be taken out as float, the session will not open. The tensor the linear operator takes in is float in both graphs, so it is safe. Actually open and run the saved version before moving on.

Measure the error at each layer boundary

Add layers <fp32> <int8> <seed> so that it feeds the same input made with the seed into the two versions and outputs the maximum absolute error, magnitude, relative error and minimum cosine at each layer boundary.

If you look only at the absolute error, the later layers always look worse, because the values themselves get larger. That is why you also output the relative error divided by the maximum absolute value of that layer's tensor. The boundary names of the two graphs differ, so pair them by position order.

Make a table of the growth ratios

Add growth <fp32> <int8> <seed> so that it outputs the relative error per layer and its ratios, whether it is monotone, the overall growth multiple and the layer with the biggest jump, and write the result for your model in /root/ep/growth.json.

If you measure, you also get stretches where the ratio is less than 1. That is because the activation function cuts off a part and the scale is set anew at each layer. That it is not monotone is itself the result, so write it as it is. If the earlier layer's relative error is 0, you cannot divide, so leave it null.

Run the same input changing only the settings

Add settings <model> <seed> so that it runs the same model and the same input changing only the optimization level and the number of threads and outputs the difference from the baseline. Write the result for your model in /root/ep/settings.json.

In this Pod, all four cases gave the same values. Write that result as it is — do not fabricate that they must differ. What matters is not the fact that they were the same but that you have a procedure for measuring whether they are the same or different. On other machines, a different answer can come out.

Write what is true only on this machine

Write /root/ep/ep_report.md in the four sections ## 무엇을 요청했고 무엇을 쥐었나 ## 그래프가 쪼개지는 자리 ## 층을 거듭하며 자란 오차 ## 이 기계에서만 참인 것 (the Korean headings mean "What was requested and what was obtained", "Where the graph splits", "The error that grew as layers stacked" and "What is true only on this machine"). The number of usable providers and the overall error growth multiple must be in it as numbers.

The last section is the core of this report. Write which of the values measured here are tied to this Pod's provider configuration and these settings. Also write that you did not measure time.