Six Nodes Down to Two: Counting What Optimization Changed
Goal
Build /root/onnxq-optimize/graph.onnx, which collects material for the optimizer to work on, and build a tool /root/onnxq-optimize/optlevel.py that runs ORT's four optimization levels, takes the graph out as a file and counts the nodes. You write down what disappears and what appears at each level, measure whether the answer is the same, and prove by a domain name why the extracted file must not be moved to another machine.
Why it matters
onnxruntime rewrites the graph when it opens a session. It folds in advance the parts computed from constants alone, removes nodes that are not needed, and groups common patterns into one. What changes and how far is decided by the optimization level.
This is often where the cause lies when the numbers do not add up in comparing speed before and after quantization. If the optimization levels of the two measurements differ, that comparison measured not quantization but optimization. If you pin the baseline to ORT_DISABLE_ALL, this wobble disappears.
A more costly incident is saving the optimized file and deploying it. That file contains nodes of a non-standard domain, so another runtime cannot open it. Yet onnx.checker passes that file — because the checker moves past unknown domains. So the basis for portability is not the checker but the list of the nodes' domains.
This lab does not measure time. This machine is an emulation, so even the same work wobbles by up to a factor of two. So it judges only how the structure changed and whether the answer is the same.
The grader does not trust the text you wrote out. It sets up a graph it built itself in a temporary directory, actually runs your tool, and checks the same file against the values the grader gets by optimizing it at the four levels. The shapes, weights and random seeds change on every run.
Steps
- Create and run /root/onnxq-optimize/build_graph.py to make /root/onnxq-optimize/graph.onnx.
- Create
nodesin /root/onnxq-optimize/optlevel.py so that it reads out the nodes and counts of the current graph. - Add
optimizeso that it takes out as a file the graph optimized at the given level. - Add
levelsso that it runs all four levels and sets the node counts side by side. - Add
fusionso that it writes down what disappeared and what appeared at each level. - Add
equalso that it feeds the same input, made with the given seed, through the four levels and measures whether the answer is the same. - Add
portableso that it judges whether the extracted file uses only standard operators. - Make a report with /root/onnxq-optimize/opt_report.json and /root/onnxq-optimize/opt_report.md.
Notes
- Python is /opt/onnx-lab/bin/python. The system
python3has neither onnx nor numpy. Example run:/opt/onnx-lab/bin/python /root/onnxq-optimize/optlevel.py levels /root/onnxq-optimize/graph.onnx - This Pod has no network. You build the materials yourself.
- Graph contract: the input
xis FLOAT with axes [symbol, 6], and the outputyis [symbol, 5]. The nodes must include MatMul, Add, Relu, Identity and Mul, and there must be at least one node whose inputs are all initializers (that is the constant to be folded). With all optimizations on, the node count must decrease. - Execution contract:
optlevel.py <명령> ...(the placeholder is the command). The answer is output as one JSON blob on standard output. On success the exit code is 0, and for an unknown command it is 2. - The level names are the four
ORT_DISABLE_ALL,ORT_ENABLE_BASIC,ORT_ENABLE_EXTENDEDandORT_ENABLE_ALL. nodes <모델>response (the placeholder is the model):{"nodes": [{"op_type", "domain", "name"}...], "count": {op_type: 개수}, "initializers": [이름 정렬]}(where the placeholders are the count and the sorted names). The order of nodes is exactly the order written in the file.optimize <모델> <단계> <출력>(the placeholders are the model, the level and the output) opens a session withSessionOptions.graph_optimization_levelset to that level andoptimized_model_filepathset to the output path, then reads the saved file again and outputs{"level", "out", "nodes": [op_type...], "count", "initializers"}.levels <모델>response (the placeholder is the model): four objects keyed by level name, each holding{"nodes": [op_type...], "count": {...}}.fusion <모델>response (the placeholder is the model): for each level,{"removed": 원본에 있었는데 사라진 op_type 정렬, "added": 새로 생긴 op_type 정렬, "total": 노드 수}(where the placeholders are the sorted op_types that were in the original but disappeared, the sorted op_types that newly appeared, and the node count).equal <모델> <씨앗> <행수>(the placeholders are the model, the seed and the number of rows) buildsnumpy.random.RandomState(씨앗).standard_normal((행수, 입력너비))(where the placeholders are the seed, the number of rows and the input width) as float32 and feeds it into all four levels. Response:{"seed", "rows", "levels": {단계: {"sum", "max", "max_abs_diff"}}, "max_abs_diff"}(where the placeholder is the level). sum and max are the sum of the whole output and the maximum absolute value, and max_abs_diff is the maximum absolute difference from the output of ORT_DISABLE_ALL. Writing to six decimal places is enough.portable <모델>(the placeholder is the model) reads the file extracted withORT_ENABLE_ALLand outputs{"nondefault_domains": 노드 도메인 가운데 표준이 아닌 것 정렬, "opset_domains": opset_import 의 도메인 정렬, "checker": "ok"|"error", "portable": 불리언}(where the placeholders are the sorted non-standard ones among the node domains, the sorted domains of opset_import, and a boolean). The standard domains are the empty string andai.onnx.- In
opt_report.json, writelevels(thenodesof the four levels),nondefault_domains,portableandequality(seed,rowsandmax_abs_diff). - Write
opt_report.mdin the four sections## 무엇이 줄었나## 어느 단계에서 무엇이 바뀌나## 답은 같은가## 이 파일을 옮겨도 되나(the Korean headings mean "What decreased", "What changes at which level", "Is the answer the same" and "Can this file be moved"), and write the node count before optimization and the remaining domain names in the text. - When extracting with
ORT_ENABLE_ALL, ORT issues a warning on standard error saying to use it only in the same environment. If you setSessionOptions.log_severity_levelto 2 you can see that warning, and with 3 it goes quiet. The warning goes to standard error, so the JSON on standard output does not break. - Floating point: if the level changes, the kernels and the order of computation change. float32 has 7 significant digits, so the last digit can wobble, so when checking for sameness, set a tolerance of about 1e-5 in relative error and write down its basis.
- Do not measure time. This machine is an emulation, so even the same work wobbles by up to a factor of two.
- Official documents: Graph optimizations · ORT Python API · Gemm · Execution Providers
- Common mistakes: deploying the optimized file as it is, using the checker passing as the basis for portability, comparing two measurements at different optimization levels, and reading the node count directly as speed.
Collect material for the optimizer to work on
Create and run /root/onnxq-optimize/build_graph.py to make /root/onnxq-optimize/graph.onnx. It must include one node whose inputs are all initializers, one Identity, and MatMul, Add, Relu and Mul.
To see constant folding, put in a node that adds two initializers and use its result later. Identity does nothing, so it is a perfect example of a node that is not needed. If you put an Add after MatMul and then a Relu, you get a pattern that gets fused. Before saving, filter it with onnx.checker.check_model(model, full_check=True).
Count the current graph
Create nodes <모델> (the placeholder is the model) in /root/onnxq-optimize/optlevel.py so that it outputs {"nodes", "count", "initializers"}. Each item of nodes holds op_type, domain and name.
The domain of a node is an empty string for a standard operator. For now they will all be empty strings, but when you open the optimized file, some will not be. To compare then, you must first count the current one. Leave the order exactly as written in the file.
Take the optimized graph out as a file
Add optimize <모델> <단계> <출력> (the placeholders are the model, the level and the output) so that it saves the graph optimized at that level to the output path, reads the saved file again and counts the nodes.
If you give a path to SessionOptions.optimized_model_filepath, the graph rewritten while opening the session is saved to that path. If you do not open a session, the file is not created either. What you extract with ORT_DISABLE_ALL is the baseline, and it is normal for the same nodes as the original to come out.
Set the four levels side by side
Add levels <모델> (the placeholder is the model) so that it runs all four levels and outputs an object holding {"nodes", "count"} for each level.
Extract one per level into a temporary directory and read them. Delete them when done. If you set the four lines side by side, you can see at a glance what happens at which level — this is faster than reading the documentation.
What disappeared and what appeared
Add fusion <모델> (the placeholder is the model) so that it outputs {"removed", "added", "total"} for each level. removed is the op_types that were in the original but disappeared, and added is the op_types that newly appeared.
Compare the op_type set of the original with the op_type set of each level. A node computed from constants alone disappears and its result settles in as an initializer. The Add after MatMul is grouped into a single Gemm, and in the next level a name that also swallowed the activation appears. Also take note of what the domain of that name is.
The graph changes, but is the answer the same
Add equal <모델> <씨앗> <행수> (the placeholders are the model, the seed and the number of rows) so that it feeds the same input, made with that seed, through all four levels and outputs {"seed", "rows", "levels", "max_abs_diff"}.
Build the input as numpy.random.RandomState(씨앗).standard_normal((행수, 입력너비)) (where the placeholders are the seed, the number of rows and the input width) in float32. Read the input width from axis 1 of the model declaration. For each level, write the sum of the whole output and the maximum absolute value, and also the maximum absolute difference from the baseline. Do not expect exactly the same; set a tolerance — float32 has only 7 significant digits.
Can this file be moved
Add portable <모델> (the placeholder is the model) so that it reads the file extracted with ORT_ENABLE_ALL and outputs {"nondefault_domains", "opset_domains", "checker", "portable"}.
The standard domains are only the empty string and ai.onnx. If there is even one node with any other domain, that file is exclusive to that environment. onnx.checker passes such a file too, so write down the checker result and the domain list together — that the two disagree is the key of this step. If you set log_severity_level to 2, you can also see the warning ORT issues while saving.
Report the four levels on one page
Write levels, nondefault_domains, portable and equality in /root/onnxq-optimize/opt_report.json, and write /root/onnxq-optimize/opt_report.md in the four sections ## 무엇이 줄었나 ## 어느 단계에서 무엇이 바뀌나 ## 답은 같은가 ## 이 파일을 옮겨도 되나 (the Korean headings mean "What decreased", "What changes at which level", "Is the answer the same" and "Can this file be moved").
The grader optimizes the same graph itself at the four levels and counts the nodes, then runs it again with the seed and number of rows written in equality and measures the difference. So measure and write the actual results. In the report, you must write the node count before optimization and the remaining domain names as numbers and names exactly so that the receiving side can judge.