TT Lab
Get started
Learn Learning paths Courses

The AI Diet Gone Wrong

What Does This File Promise? Reading What the Export Froze

Continue in TT Lab

Goal

Build an MLP yourself with onnx.helper to create /root/onnxq-export/mlp.onnx, and build a tool /root/onnxq-export/modelmeta.py that reads out the things fixed in an ONNX file. While swapping the version number, ask onnx.checker and onnxruntime in turn, directly find the spot where the checker passes but the runtime rejects, and leave it as a document.

Why it matters

An ONNX file does not hold only the graph. Which operator set version it was written with, what the IR version is, who made it, and what is an input and what is an already-decided weight are all fixed together at the moment of export. The receiving side cannot change those decisions, and to change them you have to export again. So the cause of the report "it does not open on our server" is usually not the conversion options but the first bytes of the file. Even if you lower the version number and save again, the operator definitions do not go down with it, so a file is produced that the checker passes and only the runtime rejects. The error message comes out not as "the version is low" but as "there is no implementation of this operator", so the cause is even less visible. onnx.checker is not a single layer either. The basic check looks only at the structure, and only if you give full_check=True does shape inference run. A MatMul for which the multiplication does not hold simply passes the basic check. And neither check can block an operator in an unregistered domain. The grader does not trust the text you wrote out. It sets up ONNX files it built itself in a temporary directory, actually runs your tool, and checks the same file against the answers the grader gets by reading it. The shape, names, version number and activation function change on every run.

Steps

  1. Create and run /root/onnxq-export/build_mlp.py to make /root/onnxq-export/mlp.onnx.
  2. Create info in /root/onnxq-export/modelmeta.py so that it reads out the version number, producer, inputs and outputs, initializers and nodes.
  3. Add input_overrides and runtime_inputs to info so that it reveals the boundary between initializers and inputs.
  4. Add check so that it runs onnx.checker twice, basic and with full_check=True, and writes the verdicts separately.
  5. Add load so that it writes whether onnxruntime opens the session and, if it cannot, what exception it raises.
  6. Add stamp so that it leaves the operators as they are and saves again after swapping only the version number.
  7. Add scan so that it actually measures the range of versions this runtime opens, and write your model's range in /root/onnxq-export/opset_range.json.
  8. Make a handover document with /root/onnxq-export/export_report.json and /root/onnxq-export/export_report.md.

Notes

Build the model yourself

Create and run /root/onnxq-export/build_mlp.py to make /root/onnxq-export/mlp.onnx. The input x is [symbol, 8] and the output y is [symbol, 4], and two layers are stacked with MatMul, Add and Relu. Fix the weights as initializers.

If you put a string in the shape list of helper.make_tensor_value_info, that axis becomes a symbol name, and if you put an integer, it is fixed. Make the weights with numpy_helper.from_array(배열, 이름) (the placeholders are the array and the name) and put them in the fifth argument of make_graph. Use that name as a node's input but do not put it in the input list of make_graph. Before saving, filter it once with onnx.checker.check_model(model, full_check=True).

Read out the things fixed in the file

Create info <모델> (the placeholder is the model) in /root/onnxq-export/modelmeta.py so that it outputs ir_version, producer_name, opsets, inputs, outputs, initializers and nodes as JSON.

onnx.load(path) gives you the ModelProto. model.opset_import is a list of domains and version numbers, and the default domain is an empty string. For an axis, write a symbol if there is dim_param, an integer if there is dim_value, and null if neither. Leave the initializer names out of inputs — their values are already inside the file.

Weights are not inputs

Add input_overrides (names present in both graph.input and initializer) and runtime_inputs (the input names the session actually asks for) to the info response. If the session cannot be opened, runtime_inputs is null.

Before IR 4, an initializer had to be declared in graph.input as well. That is why in files made by old tools the weights are written together in the input list, and that name means 'an input with a default value'. If you ask the runtime, you can see right away that it does not require that name as a mandatory input. Have the code that opens the session swallow the exception and return null — a file that does not open must still be readable with info.

The checker is not a single layer

Add check <모델> (the placeholder is the model) so that it runs onnx.checker in basic mode and with full_check=True and outputs {"checker", "full_check", "message"}. If it failed, put the first line of the caught exception as it is in message.

The basic check looks only at the structure. If you give full_check=True, shape inference runs too, catching a MatMul for which the multiplication does not hold and cases where the declared output shape differs from the inferred shape. Do not make up the exception text; copy the first line of str(exc) as it is — the grader checks whether the name it planted is in there.

Ask the runtime directly

Add load <모델> (the placeholder is the model) so that it tries to open an onnxruntime session and outputs {"load", "error_type", "message"}. error_type is the class name of the caught exception, and null if it opened.

Even a file that passed the checker can be rejected by the runtime. Operators of an unregistered domain, operators with no definition in that version because the version number is too low, and high versions the runtime does not yet open are like that. The kinds of rejection differ, so if you write down even the exception class name, the next person can tell the cause right away.

Just swap the version number

Add stamp <모델> <opset> <ir> <출력> (the placeholders are the model and the output) so that it changes only the opset and ir_version of the default domain and saves to a different file. The nodes, weights and producer_name must be left as they are.

Go through model.opset_import, change the version of the entry whose domain is an empty string, and add a new one if there is none. This is not a conversion but stamping — the operator definitions neither go down nor go up with it. You will see that fact with your own eyes in the next step.

Measure the range of versions this runtime opens

Add scan <모델> (the placeholder is the model) so that it swaps the version number from 1 through 27, measures whether the session opens, and outputs {"min_ok", "max_ok", "ok", "failed"}. Then write your own model's range in /root/onnxq-export/opset_range.json as min_ok and max_ok.

You just chain the stamp and load from the earlier steps together. You only need to stamp into a temporary directory and open it, so do not touch the original. The lower bound varies with from which version the operators the model uses were defined, and the upper bound varies with how far the runtime opens. The key of this step is that the two numbers have different sources.

Leave it as a one-page handover

In /root/onnxq-export/export_report.json, write the values read out of the model, the version range, and checker_ok_runtime_error holding a version that the checker passes but the runtime rejects, and write /root/onnxq-export/export_report.md in four sections.

The opset of checker_ok_runtime_error only needs to be a version lower than the lower bound that opens. Ask check and load about the file stamped with that version and write the verdicts that actually come out — the grader builds the same file and checks again. In the report, you must write the lower and upper bounds as numbers so that the receiving side can compare them with its own runtime.