Transformers — Compute Attention By Hand
Put Position In by Rotating
Goal
Build rotary position embedding (RoPE) yourself using only the standard library. Group the even dimensions into pairs, rotate each pair by the position at its own speed, and confirm in numbers that the dot product of the query and key rotated this way depends only on the difference between the two positions. Build the additive method (absolute position encoding) alongside it to compare how its values wobble at the same gap, and measure that rotation does not change the length and which components remain as the distance grows.
Why it matters
If you add a position vector, then when you expand the dot product, terms like q·P[n] and P[m]·k remain, terms containing only one of the positions. Those terms cannot be grouped into a difference, so even with exactly the same gap, the value measured near the start of a sentence differs from the value measured near the end. What matters in language is usually how many positions back the word is, yet absolute position is mixed into the score.
Rotation solves that problem in the operation itself. If you rotate two vectors together in the same direction, the angle between them does not change, so if you rotate each by its own position and then take the dot product, the value depends only on the difference. This is an equality, not an approximation, and that is why you can confirm it in numbers.
This lab does not call a model. The system Python of this Pod has no numpy and there is no internet. You build it with math alone and use only the numbers measured here. So statements like "model X uses base Y" are not made here.
The grader does not trust the explanations you wrote down. It actually imports your module, pokes at the functions with different dimensions and positions every time, and checks them against values it computes separately, within a tolerance. The inputs change on every run, so you cannot memorize values and plug them in.
Steps
- In /root/work/tf-rope/rope.py, create
DIM = 64,THETA_BASE = 10000.0,demo_vectors(),dot(a, b)andthetas(d, base=THETA_BASE).thetasreturns a different rotation speed for each pair. - Add
rotate_pair(x0, x1, angle)so that it rotates one point on the plane counterclockwise by angle radians. - Add
apply_rope(vec, pos, base=THETA_BASE)so that it groups the vector into pairs and returns a new vector in which each pair is rotated bypos * theta_i. - Create
rope_score(q, k, m, n, base=THETA_BASE)so that it rotates the query at position m and the key at position n and then takes the dot product. - Create
sin_pos(pos, d, base=THETA_BASE)andadd_score(q, k, m, n, base=THETA_BASE)so that they measure the same score with the additive method. - Create
offset_scan(q, k, offset, starts)so that it measures the same gap from several starting positions. It returns a list of(시작자리, 회전 점수, 더하는 점수)tuples (the placeholders are the starting position, the rotation score and the additive score). - Create
turns(d, delta, base=THETA_BASE)andslow_pairs(d, delta, base=THETA_BASE)so that they measure how many turns each pair makes at each gap and how many pairs have not yet gone past one turn. - Record the measured values in /root/work/tf-rope/rope_report.json and /root/work/tf-rope/rope_report.md.
Notes
- Execution contract: the grader imports
/root/work/tf-rope/rope.pyas a Python module and usesDIM,THETA_BASE,demo_vectors,dot,thetas,rotate_pair,apply_rope,rope_score,sin_pos,add_score,offset_scan,turnsandslow_pairsdirectly. It does not run it as a script, soif __name__ == "__main__"is not needed. demo_vectors()returns(q, k). They are pinned down by formula, not random numbers —q[j] = math.cos(0.7 * j + 0.3)andk[j] = math.sin(0.4 * j + 1.1), and both have lengthDIM.thetas(d, base)is a list of lengthd // 2, and the i-th value isbase ** (-2 * i / d). The first value is always 1.0 and it gets smaller toward the back. There is one per pair, not per dimension.rotate_pair(1.0, 0.0, math.pi / 2)is(0.0, 1.0). The direction is pinned to counterclockwise. The return value is a tuple of two.apply_ropegroups pairs with neighbors, as in(vec[0], vec[1]),(vec[2], vec[3]). The rotation angle of the i-th pair ispos * thetas(len(vec), base)[i], and each pair has a different angle. Do not modify the list you received in place; build a new list. Ifposis 0, the result equals the original vector.rope_score(q, k, m, n)isdot(apply_rope(q, m), apply_rope(k, n)). It does not divide by √d.sin_pos(pos, d, base)is a list of length d, and withangle = pos / (base ** ((2 * (j // 2)) / d)), the j-th value ismath.sin(angle)if j is even andmath.cos(angle)if j is odd.add_score(q, k, m, n)takes the dot product after addingsin_pos(m, len(q))to q andsin_pos(n, len(k))to k. You must not add to only one side.- For each position s in
starts,offset_scan(q, k, offset, starts)puts the query ats + offsetand the key ats. The order is the same asstarts. turns(d, delta, base)is a list of lengthd // 2, and the i-th value isdelta * theta_i / (2 * math.pi).slow_pairsis the number of pairs (an integer) whose value is less than 1.0.- The step 8 report is measured with the q and k of
demo_vectors(). The gap is 2, the starting positions are 3, 10, 100 and 4000, the positions at which the norm is measured are 0, 1, 7, 100 and 4096, and the gaps at which the turns are counted are 1, 16, 256 and 4096. The JSON keys aredim,base,offset,starts,rope_scores,rope_spread,add_scores,add_spread,norm_before,norm_max_gap,deltas,slow_pairs,turns_firstandturns_last.turns_firstis the number of turns of pair 0 at each gap, andturns_lastis the number of turns of the last pair. rope_spreadandadd_spreadare each the maximum of the four scores minus the minimum.norm_beforeis the norm of q, andnorm_max_gapis the largest absolute value among the differences between the norm after rotating at the five positions and the original norm.- Do not compare floating-point numbers with
==. The grader checks withabs(a - b) <= 1e-9 + 1e-6 * abs(b). As the position grows, the angle grows too and the error increases, so this lab judges positions only from 0 to 4096. - This Pod has no internet.
pip installdoes not work, and numpy exists only inside/opt/onnx-lab/bin/python, soimport numpydoes not work in the system Python.mathalone is enough. - Official documents: RoFormer — Rotary Position Embedding · Attention Is All You Need · Python math
- Common mistakes: assigning angles per dimension instead of per pair, using the angle of pair 0 for all pairs, reversing the rotation direction, rotating only the query and leaving the key alone, adding the position vector to only one side in the additive method, not dividing by 2π when counting turns, and modifying the list you received in place.
A different rotation speed for each pair
In /root/work/tf-rope/rope.py, create DIM = 64, THETA_BASE = 10000.0, demo_vectors(), dot(a, b) and thetas(d, base=THETA_BASE). thetas is a list of length d // 2 and its i-th value is base ** (-2 * i / d). demo_vectors() returns two lists of length DIM with q[j] = math.cos(0.7 * j + 0.3) and k[j] = math.sin(0.4 * j + 1.1).
There is one per pair, not per dimension — with 64 dimensions, that is 32. When i is 0 the exponent is 0, so the first value is 1.0, and it shrinks geometrically toward the back. If you do the division with integers, every exponent becomes 0 and all the values are 1.0, so use floating point, like -2.0 * i / d. dot is one line: pair up with zip, multiply, and add.
Rotate once on the plane
Add rotate_pair(x0, x1, angle). It returns a tuple of two, the point (x0, x1) rotated counterclockwise by angle radians. rotate_pair(1.0, 0.0, math.pi / 2) is (0.0, 1.0).
It is (x0*cos - x1*sin, x0*sin + x1*cos). Flipping just one of the two signs makes it clockwise, and then all the values in the later steps change. If the angle is 0 it must give back the original point, and for any angle the distance from the origin must not change — because the squares of cos and sin add up to 1.
Rotate a vector by its position
Add apply_rope(vec, pos, base=THETA_BASE). Group pairs with neighbors, as in (vec[0], vec[1]), (vec[2], vec[3]), and return a new list in which the i-th pair is rotated by pos * thetas(len(vec), base)[i]. Leave the list you received as it is.
Each pair has a different angle — if you use the angle of pair 0 for all pairs, it is just rotating the whole thing, and only a single layer of position information goes in. If pos is 0, all the angles are 0 and the result must equal the original vector, and for any position the norm of the vector must not change. Do not forget to take base and pass it to thetas — if you use only the default, it is silently wrong when called with a different base.
The score of a rotated query and key
Create rope_score(q, k, m, n, base=THETA_BASE). It returns the floating-point number obtained by rotating the query at position m and the key at position n each and then taking the dot product. It does not divide by √d.
It is one line: dot(apply_rope(q, m, base), apply_rope(k, n, base)). If you rotate only one side, absolute position remains as it is and the property in the later steps collapses. When m and n are equal, you must get the same value as the unrotated dot(q, k) — because both were rotated together in the same direction, so the angle between them stays the same.
Put the additive method alongside
Create sin_pos(pos, d, base=THETA_BASE) and add_score(q, k, m, n, base=THETA_BASE). sin_pos is a position vector of length d, and with angle = pos / (base ** ((2 * (j // 2)) / d)), its j-th value is math.sin(angle) if j is even and math.cos(angle) if j is odd. add_score takes the dot product after adding a position vector to each of q and k.
This side is the method of building values and adding them. If you add to only one side, the comparison does not hold. When you expand the dot product, besides q·k you get q·P[n] and P[m]·k, and since these two terms contain only one of the positions, they cannot be grouped into a difference — you will see the result in numbers in the next step.
The same gap gives the same score
Create offset_scan(q, k, offset, starts). For each position s in starts, put the query at s + offset and the key at s and measure the score with both methods. The return value is a list of (s, 회전 점수, 더하는 점수) tuples (the placeholders are the rotation score and the additive score), in the same order as starts. After building it, print offset_scan(q, k, 2, [3, 10, 100, 4000]) yourself and confirm with your own eyes that the four rotation values are the same while the additive side wobbles.
If you call the rope_score and add_score you built earlier as they are, it is five lines. If you swap the query and key, the sign of the gap flips and you get a different value. The four rotation values should agree to twelve decimal places — they are not exactly the same but a very small difference remains, and that is floating-point error. Compare the scatter on the additive side against that magnitude.
Which pairs remain as it gets farther
Create turns(d, delta, base=THETA_BASE) and slow_pairs(d, delta, base=THETA_BASE). turns is a list of length d // 2, and its i-th value is delta * theta_i / (2 * math.pi), that is, the number of turns that pair makes. slow_pairs is the number of pairs (an integer) whose value is less than 1.0.
One turn is 2π radians. If you forget to divide, you get the angle instead of the number of turns and the counts come out completely different. A pair that has gone past one turn writes that distance and the distance with one turn subtracted as the same angle, so it cannot tell the two apart. Call it with gaps 1, 16, 256 and 4096 and see how the number of remaining pairs decreases.
Leave what you measured as a record
Measure with the q and k of demo_vectors(), write dim, base, offset, starts, rope_scores, rope_spread, add_scores, add_spread, norm_before, norm_max_gap, deltas, slow_pairs, turns_first and turns_last in /root/work/tf-rope/rope_report.json, and write /root/work/tf-rope/rope_report.md in the four sections ## 무엇을 쟀나 ## 같은 간격이면 같은 점수다 ## 더하는 방식은 왜 다른가 ## 멀어지면 어느 성분이 남는가 (the Korean headings mean "What was measured", "The same gap gives the same score", "Why the additive method differs" and "Which component remains as it gets farther"). The gap is 2, the starting positions are 3, 10, 100 and 4000, the positions at which the norm is measured are 0, 1, 7, 100 and 4096, and the gaps at which the turns are counted are 1, 16, 256 and 4096.
Do not write the numbers by hand; fill them in with values obtained by actually running your own code. Take rope_scores and add_scores from the tuples offset_scan returned. rope_spread and add_spread are each the maximum minus the minimum, and norm_max_gap is the largest absolute value of the differences measured at the five positions. turns_first is the number of turns of pair 0 at each gap, and turns_last is the number of turns of the last pair. In the body of the report, write in numbers the first rotation score, the scatter on the additive side and the norm of q — the grader checks whether those three values are in the text.