TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

What Freezes at Export Time: Opset, IR, Initializers

Continue in TT Lab

In one line

An ONNX file does not hold only the graph; it also fixes which opset version it was written with, which IR version it was written in, and what is an input and what is an already-decided value. These four are decided at the moment of export, and the receiving side cannot change those decisions.

Why this was needed

A report that quantization failed usually does not start from the conversion command. It starts with "this file does not open on our server". The file looks fine, and onnx.checker passes it without a word. Yet the runtime stops at opening the session.

In a situation like this, what eats the most time is looking for the cause outside the file. You change the conversion options, raise the runtime version, and suspect the hardware. In fact the answer is written in the first few bytes of the file.

Four things that are fixed in the file

First, the operator set version and domain. The model writes a version number for each domain in opset_import. If the domain is an empty string, it is the standard operators (ai.onnx), and another domain such as com.microsoft is an extension of a specific runtime. The version number is an instruction "read the operators in this file by the definitions of this version", not a performance setting.

Second, the IR version. It is the format version of the container that holds the graph. ONNX Versioning numbers opset and IR separately for this reason — the speed at which operator definitions change and the speed at which the file format changes are different. In fact, before IR 4, an initializer also had to be declared in graph.input, and after that it does not have to. That is why, in files made by old tools, the weights are still written together in the input list.

Third, producer information. producer_name and producer_version record who made this file. They do no work, but they are the first lines you look at when an incident occurs. If a file with them empty arrives, that itself is the first problem.

Fourth, the boundary between initializers and inputs. Weights go into graph.initializer, and values a person supplies at execution time stay in graph.input. If a name is written in both, it means "an input with a default value", and the runtime does not require that name as a mandatory input. So if you count the number of inputs by looking only at the file, you get it wrong — ONNX Concepts explains the two separately.

What the checker catches and what it does not

onnx.checker.check_model(model) by default looks at the structure. Whether the nodes are in topological order, whether the names they point to exist, whether the required fields are there. If you give it full_check=True, it runs even shape inference. This difference is large in practice. A MatMul with shapes for which the multiplication does not hold passes the basic check as it is and gets caught only in full_check.

And there are things neither check catches. An operator in an unregistered domain is passed by the checker — because the checker sees an unknown domain as "someone's extension" and moves on. A file stamped with a lowered version number passes in the same way. In both cases it stops at the runtime.

onnx.checker(기본)      구조만
onnx.checker(full)      구조 + 모양 추론
onnxruntime 세션 열기    구조 + 모양 + 이 런타임에 그 커널이 있는가

The third line is the narrowest, and what deployment actually requires is the third line.

What it looks like in the field

First, "please lower the version". It is common to lower only the opset number and save again because the receiving runtime is old. But the version number is only a stamp, and the operator definitions do not go down with it. If the lowered version has no definition of that operator, the checker passes it and the runtime rejects it. The error message also comes out not as "the version is low" but as "cannot find an implementation of this operator", so the cause is not visible.

Second, the versions the exporting side can write are wider than the versions the running side opens. The library can write in the latest version, but the runtime does not yet open that version. This is why ONNX Runtime Compatibility lists in a table the opset range each runtime version opens. So the advice "export with the latest version" is itself dangerous.

Third, weights look like inputs. When you open a file made by an old tool, it looks like five inputs, but the runtime asks for only one. If you write code that fills in five inputs without knowing this difference, that code pushes the weights in anew on every call.

Fourth, using the checker passing as the basis for deployment. Many teams put "checker passes" in the release conditions. That condition is necessary but not sufficient. The deployment gate must include a step that actually opens a session with the same runtime version as the receiving side.

What really matters in practice

What you will do in the next lab

You build a two-layer MLP yourself with onnx.helper, and grow one step at a time a tool modelmeta.py that reads out the things fixed in that file. You meet an old-format file that also declares initializers as inputs and separate inputs from weights, measure directly the difference between the basic check and full_check, and, swapping only the version number, find the lower and upper bounds of the range the runtime opens. The grader builds its own files each time with a different shape, name, version number and activation function, actually runs your tool, and checks the answers against its own.