One Frozen Axis Blocks the Whole Batch
In one line
An axis in ONNX is one of three: fixed as an integer, open as a symbol name, or absent altogether, and most incidents where you cannot raise the batch come not from the third but from the first — and they start at a spot that passes without any error.
Why this was needed
To measure how much faster a quantized model got, you have to raise the batch. You tried to put in 32 at once, and the runtime rejects it. You were clearly told the model was exported with a "dynamic batch", and the session opened fine. But it stops the moment you put in values.
The common response here is to dig through runtime options. The answer is inside the file. A session opening only means the graph holds together; it does not mean it accepts the shape we are about to put in.
An axis has three states
The shape of a tensor as described in ONNX Concepts is, for each axis, one of three.
- Integer:
dim_valueis set. That axis must be that value. - Symbol name:
dim_paramholds a string. It is decided at execution time, and the same name means the same value. - Empty: neither is there. It means unknown, and inference stops here too.
The "same name means the same value" of the second is often overlooked. If two inputs both have batch written on axis 0, that is not simply "both dynamic" but a constraint that the two must have the same number of rows. If you put 5 rows in one and 3 rows in the other, it is rejected.
But that rejection is not friendly. The runtime does not tell you "the symbol batch split into 5 and 3". It cites the name of a node after graph optimization and says the input shape of that node does not match. There is a layer of optimization between the cause and the symptom, so at first glance the cause is not visible.
선언 x [batch, 4] bias [batch, 3]
넣은 것 x 5행 bias 3행
런타임의 말 (융합된 노드 이름) 의 입력 모양이 맞지 않는다
실제 원인 같은 심볼에 다른 값을 넣었다
What inference fills and what it cannot
onnx.shape_inference sweeps the graph, fills in the shapes of intermediate tensors and puts them in value_info. If the input is [batch, 4] and the weight is 4 x 6, the intermediate tensor is filled as [batch, 6]. A symbol propagates as a symbol.
There are two cases it cannot fill. One is when the shape of the input is empty. If you start from the unknown, you stay unknown to the end. The other is when there is no intermediate tensor at all. If there is only one node and its output is the graph output, there is nothing to fill, so value_info comes back empty. You must not read this as "inference failed" — there was just nothing to fill.
And there is a more dangerous case. Inference filling in a wrong value with confidence. If the target shape of a Reshape is pinned as a constant, inference trusts that constant as it is. Even if you left the input open as [batch, 4], if Reshape points at [3, 6], every tensor after that is fixed at 3. No error occurs. The batch axis remains only in the declaration and is actually dead.
What it looks like in the field
First, a model exported with batch 1. If you give the export tool one example input, that shape is fixed as it is. If you do not specify dynamic axes, [1, ...] is pinned, and that file handles only one item at a time forever. An experiment to measure throughput becomes entirely meaningless.
Second, the rejection comes not at opening the session but at putting in values. The session opens fine. So you conclude "the model loaded, so it is not a model problem" and dig in the wrong place. The error text for a fixed axis states exactly which axis of which input expected what and received what, so you only need to read that line.
Third, a Reshape with a constant shape pinned. This is the quietest incident. If you open the model, the input is open fine as [batch, 4], and the inference result is consistent too. It is just that the batch axis has turned into an integer from some point on. That is why "I exported with a dynamic axis, so why doesn't it work" keeps coming up. The fix is simple — if you put -1 in the batch position, it is computed from the other axes and the axis comes back to life.
Fourth, quantization does not fix axes, but it hides the cause. A quantized graph has its node names changed and Q/DQ inserted, which makes it hard to read. So checking axis problems before quantization is much cheaper.
What really matters in practice
- Try it by putting values in and then speak. Do not just read the declarations; leave a record of actually putting in several numbers of rows.
- The same symbol is the same value. If two inputs use the same name, write that constraint in the documentation.
- Distinguish inference's blanks from its confidence. Being empty and being wrongly filled are entirely different problems.
- Use -1 for the batch position. Do not write the batch value directly in the target shape of Reshape.
What you will do in the next lab
You build and set side by side a version with the batch axis as a symbol and a version fixed to 1 using the same weights, and grow the tool axes.py one step at a time. You put several numbers of rows into the two files with the same code and record where they diverge, separate what inference filled from what it could not, and try putting different numbers of rows into two inputs that use the same symbol. Finally you find a Reshape with a constant shape pinned and restore the batch position to bring the axis back to life. The grader builds its own files each time with different axis names and shapes and different pinned row counts, actually runs your tool, and checks the answers against its own.