Transformers — Compute Attention By Hand
Rotate the Position Instead of Adding It
In one line
If instead of adding the position to the vector you rotate the vector by the position, the dot product of the query and the key depends only on the difference between the two positions. The value measured at positions 3 and 5 and the value measured at positions 4000 and 4002 become the same.
Why this was needed
Earlier we built sine/cosine position vectors and added them to the input. That clearly put order in. The problem is how it went in.
The attention score is decided by a single dot product of the query and the key. If you add a position vector P and expand the dot product, four terms come out.
(q + P[m]) · (k + P[n])
= q·k + q·P[n] + P[m]·k + P[m]·P[n]
The first term is the score between contents, and the last term is the score between positions. The problem is the two middle terms. q·P[n] contains only n, and P[m]·k contains only m. A term containing only one of the positions cannot be grouped into a difference. So even when the gap is exactly two, the value measured near the start of a sentence differs from the value measured near the end.
Why is this bad? What matters in language is usually "how many positions back is the word", not "what number character of the document is this". The noun that a modifying clause describes is right after it, and what a pronoun refers to is a few sentences earlier. These are all relationships written as distance. Yet the score the model sees has absolute position mixed into it. If a position number never seen in training comes up in service — and that is exactly what happens the moment you feed in a longer text — no one knows what value those terms will take.
There have been attempts to fix this for a long time. If you train the position vectors, they simply stop at positions not in the table, and if you add a constant that differs per distance to the score, it only adds one more term and the problem of content and position being mixed remains. Neither touched the dot product itself.
So the question posed by the RoFormer paper is this. Can we make the dot product see only relative position from the start? Not by fixing it after the score is produced, but so that it is already that way at the point where the query and key are made.
Rotate instead of adding
The answer is to change the operation. Do not add; rotate.
Group the even dimensions two by two and see them as points on a plane. With 64 dimensions there are 32 pairs. Fix an angle for each pair, and the vector at position m rotates each pair by m times that angle.
theta_i = base ** (-2*i/d) # i 번째 쌍의 회전 속도
(x0, x1) -> (x0*cos(a) - x1*sin(a), x0*sin(a) + x1*cos(a)) # a = m * theta_i
Pair 0 turns fastest at speed 1, and the speeds slow down geometrically toward the back. The base is 10000 in the paper's definition.
What matters here is that nothing was added. No new value was mixed into the vector. Only the direction of the existing values was turned by the position.
It is also worth noting that the rotation happens only within a pair. Pair 0 mixes only dimensions 0 and 1, and pair 1 mixes only dimensions 2 and 3. It is not a big matrix that shuffles all the dimensions; it is small 2 by 2 rotations lined up along the diagonal. So instead of doing one whole matrix multiplication, it ends with multiplying by cos and sin for each pair and adding.
Why only relative position remains
Looking at just one pair, it fits on one line. The transpose of a rotation matrix R is the rotation in the opposite direction.
(R(m*theta) q) · (R(n*theta) k) = q · R(m*theta)ᵀ R(n*theta) k
= q · R((n-m)*theta) k
Where m and n each are disappears, and only the difference remains. This happens separately for each pair, and since the dot product adds them up, the same statement holds for the whole vector.
In words, it is like this. If you turn two things together in the same direction, the angle between them does not change. It is like turning two clock hands as a whole and the angle between them stays the same. A dot product is ultimately determined by the lengths and the angle between, so if the angle between does not change, the score does not change.
And the rotation does not change the length. That is because the squares of cos and sin add up to 1. Adding puts another value on top of the original signal and changes its magnitude, but rotation puts nothing on top. It puts in position information while leaving the content untouched.
What remains as it gets farther
This is where the different speed of each pair pays off.
Between two positions separated by a gap delta, the i-th pair opens by delta * theta_i. Dividing this by one full turn (2π) gives how many turns it has made. A fast pair makes several turns even with a small separation. A pair that has gone past one turn writes delta and "the distance with one turn subtracted from delta" as the same angle — meaning you cannot tell the two apart by looking at that pair alone.
So as the distance grows, only the slow pairs that have not yet made one turn tell that distance apart properly. Short distances are divided finely by the fast pairs, and long distances are divided coarsely by the slow pairs. This is why the speeds are laid out geometrically. It is like stacking several rulers: the ruler with fine marks measures only short things, and the ruler with coarse marks measures long things.
If you measure it yourself with 64 dimensions in the lab, the numbers come out clearly. When the gap is 1, all 32 pairs are within their first turn, but as the gap grows, that number shrinks. Once you have seen the shape of the shrinking with your own eyes, you can see why "extending the context" leads to a story about adjusting angles — because you have to make the remaining rulers longer, or re-lay the rulers more sparsely.
What it looks like in the field
First, you extend the context and quality collapses. Once you go past the length used in training, even the slow pairs enter at angles never seen before. The angle itself is computed, but the model has never learned what to do at that angle.
Second, people get confused about which layer to put it in. The approach of adding a position vector ends with adding it once to the input embedding. Rotation is not like that — it is applied at the point where attention uses the query and key. The value (V) is not rotated. If you rotate it, the content gets dragged around by position.
Third, the query and key use different conventions. An implementation that groups pairs by neighbors and an implementation that pairs the first half with the second half are both common. The properties are the same either way, but putting the weights of one into the code of the other is silently wrong. No error occurs; only the scores become strange.
Fourth, you put rotated values in the cache and then count the positions again. If you keep already-rotated keys in the cache and rotate them again later, the position goes in twice. It drifts more the longer the generation gets.
Fifth, changing the base makes it a different model. The base is the value that decides the speeds of all the pairs at once. If you run inference with a base different from the trained one, the angles of every pair go off.
What really matters in practice
- Pin down with a test that the same gap gives the same score. Measuring while moving the position and checking that the values are equal is a test of a few lines, but if it breaks, everything below loses its meaning.
- Rotation applies only to the query and key. It does not go into the value.
- Check that length is preserved. If the norm changed after rotating, you did something other than a rotation.
- Be prepared for floating-point error. As the position grows, the angle grows too, and the error increases. When checking equality, use a comparison with a tolerance rather than
==.
What you will do in the next lab
You grow /root/work/tf-rope/rope.py one step at a time. You use neither numpy nor torch — the system Python of this Pod has no numpy (it exists only inside /opt/onnx-lab/bin/python) and there is no internet either. The standard library math alone is enough.
You start by making a different rotation speed for each pair and build up through one 2-dimensional rotation, rotating a whole vector by its position, and the score of a rotated query and key. Then you build the additive method alongside it and compare the two values at the same positions.
The heart of it is step six. With the gap fixed at two, you move the starting position through 3, 10, 100 and 4000 and measure the score with both methods. The rotation side gives the same value at all four positions, and the additive side wobbles. You will see the width of that scatter in numbers.
Finally, as the gap grows, you count how many turns each pair makes, and measure how many pairs remain that have not yet gone past one turn. The grader does not trust the explanations you wrote down — it actually imports your module, pokes at the functions with different dimensions and positions every time, and checks them against values it computes separately, within a tolerance.