TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

The Providers You Ask For Are Not the Ones You Get

Continue in TT Lab

In one line

An execution provider (EP) is something you request, not something you are guaranteed. Even if you put in an EP that does not exist, it is not an error but a silent fallback, and you can know what you actually got only by asking the session.

Why this becomes a problem

It often happens that you wrote in the list that you would use an accelerator but it is actually running on the CPU. The problem is that nothing happens at that time. No exception, and the exit code is 0. Only one line of warning passes through standard error, and in a place that gathers logs, that one line gets buried.

If you measure directly in this lab environment, it goes like this. Even if you put a nonexistent CUDAExecutionProvider at the very front of the list, the session is created and inference works. If you ask session.get_providers(), only CPUExecutionProvider comes back. If you make up and put in a name that does not exist at all, a message "Unknown Provider Type" appears, and it likewise falls to the CPU and runs.

So there is one rule. Do not trust the request list; ask the session. What you must leave in the deployment log is not the list you requested but the list the session returned.

What an EP decides

An EP is not a device that chooses "where to run this model", but a device that decides who takes charge of each node. The executor sweeps the request list in order, asks each EP "can you take this node?", and takes only the nodes it says it will take. The remaining nodes go to the next EP, and if finally no one takes them, the CPU does.

So a single model splits into several pieces that run in different places. Between pieces, values must be handed over. If the list of supported operators keeps getting cut along the graph order, the number of pieces grows, and the number of handoffs grows too.

지원: MatMul 만               조각 5개 · 넘김 4번
  [MatMul] -> Add -> Relu -> [MatMul] -> Add
    EP        CPU    CPU        EP       CPU

The order of the list is the priority. The provider written first asks first and takes the nodes it says it will take. So even if you write the same two providers with only the order changed, the placement can differ. And the CPU follows at the end even if you do not write it — in this environment, if you write only AzureExecutionProvider and create a session, the actual list becomes that and CPUExecutionProvider, two. That is because you cannot leave a node with no one to take it. This is also why the requested list and the actual list differing is normal behavior, not an error.

An important conclusion comes out here. Even if you feed the same input to the same model, different machines can produce different numbers. If the pieces split differently, the order of computation changes, and with floating-point addition, if the order changes, the result differs subtly. So "it was right on our laptop" is hard to use as evidence.

What happens to the error as you stack layers

Quantization error newly arises at each layer, and the error that arose in an earlier layer goes into the next layer's input. So the error accumulates. But if you actually measure it, it does not grow evenly per layer. When this lab measured a five-layer model, the relative error at the layer boundaries was 0.099, 0.074, 0.063, 0.214, 0.167 and 0.369. It went down and up, and overall became 3.7 times.

The reason there are stretches where it decreases is that the magnitude of values differs from layer to layer, the activation function cuts off a part, and the scale is set anew. So you must not end with the story "it grows a little at each layer"; you must measure at each layer boundary and make a table to see where the problem is. If the ratio jumps threefold at some layer, that layer is a candidate for narrowing the scope.

Even if the error grows, accuracy does not collapse by that much. Even if the relative error at the last boundary is 0.37, the number of samples whose decision changes can be few. Because the output vector as a whole being pushed a little and the ranking being flipped are different matters. So use the per-layer error table to choose where to work on, and judge whether to deploy separately with the task metric. If you try to merge the two into one number, both get blurred.

To measure layer boundaries, you have to be able to see the intermediate tensors. In an ONNX graph, you can just append the tensors you want to see to the graph output list. But in an integer version you cannot take out an integer tensor as float, so it is safe to take the side the linear operator takes in as the boundary. That spot is float in both graphs.

What you have to do in the field

What you will do in the next lab

You request a mix of existing and nonexistent EPs and measure directly what you end up with, and build a simulation that counts how many pieces the graph splits into when a list of supported operators is given. Then you take the layer-boundary tensors out as graph outputs, run the FP32 version and the integer version on the same input to measure the error at each layer, and make a table of the growth ratios. Finally you run the same model and same input changing only the optimization level and the number of threads to see whether the result changes. The grader actually runs your tool each time with a different model, seed and support list and redoes the same computation to check against it.