What Does This File Promise? Reading What the Export Froze
Goal
Build an MLP yourself with onnx.helper to create /root/onnxq-export/mlp.onnx, and build a tool /root/onnxq-export/modelmeta.py that reads out the things fixed in an ONNX file. While swapping the version number, ask onnx.checker and onnxruntime in turn, directly find the spot where the checker passes but the runtime rejects, and leave it as a document.
Why it matters
An ONNX file does not hold only the graph. Which operator set version it was written with, what the IR version is, who made it, and what is an input and what is an already-decided weight are all fixed together at the moment of export. The receiving side cannot change those decisions, and to change them you have to export again.
So the cause of the report "it does not open on our server" is usually not the conversion options but the first bytes of the file. Even if you lower the version number and save again, the operator definitions do not go down with it, so a file is produced that the checker passes and only the runtime rejects. The error message comes out not as "the version is low" but as "there is no implementation of this operator", so the cause is even less visible.
onnx.checker is not a single layer either. The basic check looks only at the structure, and only if you give full_check=True does shape inference run. A MatMul for which the multiplication does not hold simply passes the basic check. And neither check can block an operator in an unregistered domain.
The grader does not trust the text you wrote out. It sets up ONNX files it built itself in a temporary directory, actually runs your tool, and checks the same file against the answers the grader gets by reading it. The shape, names, version number and activation function change on every run.
Steps
- Create and run /root/onnxq-export/build_mlp.py to make /root/onnxq-export/mlp.onnx.
- Create
infoin /root/onnxq-export/modelmeta.py so that it reads out the version number, producer, inputs and outputs, initializers and nodes. - Add
input_overridesandruntime_inputstoinfoso that it reveals the boundary between initializers and inputs. - Add
checkso that it runsonnx.checkertwice, basic and withfull_check=True, and writes the verdicts separately. - Add
loadso that it writes whether onnxruntime opens the session and, if it cannot, what exception it raises. - Add
stampso that it leaves the operators as they are and saves again after swapping only the version number. - Add
scanso that it actually measures the range of versions this runtime opens, and write your model's range in /root/onnxq-export/opset_range.json. - Make a handover document with /root/onnxq-export/export_report.json and /root/onnxq-export/export_report.md.
Notes
- Python is /opt/onnx-lab/bin/python. The system
python3has neither onnx nor numpy. Example run:/opt/onnx-lab/bin/python /root/onnxq-export/modelmeta.py info /root/onnxq-export/mlp.onnx - This Pod has no network. Installation does not work, and there is no model already downloaded. You build the materials yourself.
- Model contract: input
xis FLOAT with 2 axes, axis 0 is a symbol name (a string) and axis 1 is 8. Outputyis FLOAT with axis 1 being 4. The nodes include MatMul, Add and Relu, and the weights are fixed as at least 2 initializers.producer_nameis not left empty. The opset is between 7 and 26, and ir_version is at most 13. - Execution contract:
modelmeta.py <명령> ...(the placeholder is the command). The answer is output as one JSON blob on standard output. On success the exit code is 0, and for an unknown command it is 2. Warnings that onnxruntime prints to standard error are not the answer, so keep only standard output clean. info <모델>response (the placeholder is the model):ir_versionis an integer,producer_nameis a string,opsetsis an object keyed by domain,inputsandoutputsare lists of{"name", "elem_type", "dims"},initializersis a list with names sorted, andnodesis a list of op_type. Each axis ofdimsis an integer if fixed, a string if a symbol, and null if there is nothing.elem_typeis the name given byonnx.TensorProto.DataType.Name(...)(for example FLOAT).- From step 3,
input_overrides(a sorted list of the names present in both graph.input and initializer) andruntime_inputs(the input names the session actually asks for, or null if the session cannot be opened) are added to theinforesponse. Names filled by initializers are not put ininputs. check <모델>response (the placeholder is the model):{"checker": "ok"|"error", "full_check": "ok"|"error", "message": 문자열}(where the placeholder is a string). When it fails, message carries the first line of the caught exception as it is.load <모델>response (the placeholder is the model):{"load": "ok"|"error", "error_type": 예외 클래스 이름 또는 null, "message": 문자열}(where the placeholders are the exception class name or null, and a string).stamp <모델> <opset> <ir> <출력>(the placeholders are the model and the output) changes only the opset and ir_version of the default domain and saves to a different file. It does not touch the nodes, weights or producer_name. The response is{"out", "opset", "ir_version", "nodes"}.scan <모델>response (the placeholder is the model):{"min_ok": 정수 또는 null, "max_ok": 정수 또는 null, "ok": 정수 목록, "failed": 정수 목록}(where the placeholders are an integer or null, and lists of integers). It swaps the version number from 1 through 27 and looks only at whether the session opens (it judges by opening the session, not by the checker result).- In
opset_range.json, write at leastmin_okandmax_ok. Inexport_report.json, writemodel,ir_version,producer_name,opset,nodes,initializers,runtime_inputs,min_ok_opsetandmax_ok_opset, and achecker_ok_runtime_errorobject (opset,checkerandload) holding a version that the checker passes but the runtime rejects. - Write
export_report.mdin the four sections## 무엇을 내보냈나## 판이 굳는 자리## 검사기가 못 잡는 것## 다음 사람에게 넘길 것(the Korean headings mean "What was exported", "Where the version gets fixed", "What the checker does not catch" and "What to hand over to the next person"), and write the lower and upper bounds of the versions that open as numbers. - Official documents: ONNX Concepts · ONNX Versioning · ONNX IR · ORT Compatibility · ORT Python API
- Common mistakes: counting
graph.inputand calling it the number of inputs, running only the basic check and reporting a pass, believing that lowering the version number lowers the operators too, and mixing runtime warnings into standard output and breaking the JSON.
Build the model yourself
Create and run /root/onnxq-export/build_mlp.py to make /root/onnxq-export/mlp.onnx. The input x is [symbol, 8] and the output y is [symbol, 4], and two layers are stacked with MatMul, Add and Relu. Fix the weights as initializers.
If you put a string in the shape list of helper.make_tensor_value_info, that axis becomes a symbol name, and if you put an integer, it is fixed. Make the weights with numpy_helper.from_array(배열, 이름) (the placeholders are the array and the name) and put them in the fifth argument of make_graph. Use that name as a node's input but do not put it in the input list of make_graph. Before saving, filter it once with onnx.checker.check_model(model, full_check=True).
Read out the things fixed in the file
Create info <모델> (the placeholder is the model) in /root/onnxq-export/modelmeta.py so that it outputs ir_version, producer_name, opsets, inputs, outputs, initializers and nodes as JSON.
onnx.load(path) gives you the ModelProto. model.opset_import is a list of domains and version numbers, and the default domain is an empty string. For an axis, write a symbol if there is dim_param, an integer if there is dim_value, and null if neither. Leave the initializer names out of inputs — their values are already inside the file.
Weights are not inputs
Add input_overrides (names present in both graph.input and initializer) and runtime_inputs (the input names the session actually asks for) to the info response. If the session cannot be opened, runtime_inputs is null.
Before IR 4, an initializer had to be declared in graph.input as well. That is why in files made by old tools the weights are written together in the input list, and that name means 'an input with a default value'. If you ask the runtime, you can see right away that it does not require that name as a mandatory input. Have the code that opens the session swallow the exception and return null — a file that does not open must still be readable with info.
The checker is not a single layer
Add check <모델> (the placeholder is the model) so that it runs onnx.checker in basic mode and with full_check=True and outputs {"checker", "full_check", "message"}. If it failed, put the first line of the caught exception as it is in message.
The basic check looks only at the structure. If you give full_check=True, shape inference runs too, catching a MatMul for which the multiplication does not hold and cases where the declared output shape differs from the inferred shape. Do not make up the exception text; copy the first line of str(exc) as it is — the grader checks whether the name it planted is in there.
Ask the runtime directly
Add load <모델> (the placeholder is the model) so that it tries to open an onnxruntime session and outputs {"load", "error_type", "message"}. error_type is the class name of the caught exception, and null if it opened.
Even a file that passed the checker can be rejected by the runtime. Operators of an unregistered domain, operators with no definition in that version because the version number is too low, and high versions the runtime does not yet open are like that. The kinds of rejection differ, so if you write down even the exception class name, the next person can tell the cause right away.
Just swap the version number
Add stamp <모델> <opset> <ir> <출력> (the placeholders are the model and the output) so that it changes only the opset and ir_version of the default domain and saves to a different file. The nodes, weights and producer_name must be left as they are.
Go through model.opset_import, change the version of the entry whose domain is an empty string, and add a new one if there is none. This is not a conversion but stamping — the operator definitions neither go down nor go up with it. You will see that fact with your own eyes in the next step.
Measure the range of versions this runtime opens
Add scan <모델> (the placeholder is the model) so that it swaps the version number from 1 through 27, measures whether the session opens, and outputs {"min_ok", "max_ok", "ok", "failed"}. Then write your own model's range in /root/onnxq-export/opset_range.json as min_ok and max_ok.
You just chain the stamp and load from the earlier steps together. You only need to stamp into a temporary directory and open it, so do not touch the original. The lower bound varies with from which version the operators the model uses were defined, and the upper bound varies with how far the runtime opens. The key of this step is that the two numbers have different sources.
Leave it as a one-page handover
In /root/onnxq-export/export_report.json, write the values read out of the model, the version range, and checker_ok_runtime_error holding a version that the checker passes but the runtime rejects, and write /root/onnxq-export/export_report.md in four sections.
The opset of checker_ok_runtime_error only needs to be a version lower than the lower bound that opens. Ask check and load about the file stamped with that version and write the verdicts that actually come out — the grader builds the same file and checks again. In the report, you must write the lower and upper bounds as numbers so that the receiving side can compare them with its own runtime.