TT Lab
Get started
Learn Learning paths Courses

Transformers — Compute Attention By Hand

Put Position In by Rotating

Continue in TT Lab

Goal

Build rotary position embedding (RoPE) yourself using only the standard library. Group the even dimensions into pairs, rotate each pair by the position at its own speed, and confirm in numbers that the dot product of the query and key rotated this way depends only on the difference between the two positions. Build the additive method (absolute position encoding) alongside it to compare how its values wobble at the same gap, and measure that rotation does not change the length and which components remain as the distance grows.

Why it matters

If you add a position vector, then when you expand the dot product, terms like q·P[n] and P[m]·k remain, terms containing only one of the positions. Those terms cannot be grouped into a difference, so even with exactly the same gap, the value measured near the start of a sentence differs from the value measured near the end. What matters in language is usually how many positions back the word is, yet absolute position is mixed into the score. Rotation solves that problem in the operation itself. If you rotate two vectors together in the same direction, the angle between them does not change, so if you rotate each by its own position and then take the dot product, the value depends only on the difference. This is an equality, not an approximation, and that is why you can confirm it in numbers. This lab does not call a model. The system Python of this Pod has no numpy and there is no internet. You build it with math alone and use only the numbers measured here. So statements like "model X uses base Y" are not made here. The grader does not trust the explanations you wrote down. It actually imports your module, pokes at the functions with different dimensions and positions every time, and checks them against values it computes separately, within a tolerance. The inputs change on every run, so you cannot memorize values and plug them in.

Steps

  1. In /root/work/tf-rope/rope.py, create DIM = 64, THETA_BASE = 10000.0, demo_vectors(), dot(a, b) and thetas(d, base=THETA_BASE). thetas returns a different rotation speed for each pair.
  2. Add rotate_pair(x0, x1, angle) so that it rotates one point on the plane counterclockwise by angle radians.
  3. Add apply_rope(vec, pos, base=THETA_BASE) so that it groups the vector into pairs and returns a new vector in which each pair is rotated by pos * theta_i.
  4. Create rope_score(q, k, m, n, base=THETA_BASE) so that it rotates the query at position m and the key at position n and then takes the dot product.
  5. Create sin_pos(pos, d, base=THETA_BASE) and add_score(q, k, m, n, base=THETA_BASE) so that they measure the same score with the additive method.
  6. Create offset_scan(q, k, offset, starts) so that it measures the same gap from several starting positions. It returns a list of (시작자리, 회전 점수, 더하는 점수) tuples (the placeholders are the starting position, the rotation score and the additive score).
  7. Create turns(d, delta, base=THETA_BASE) and slow_pairs(d, delta, base=THETA_BASE) so that they measure how many turns each pair makes at each gap and how many pairs have not yet gone past one turn.
  8. Record the measured values in /root/work/tf-rope/rope_report.json and /root/work/tf-rope/rope_report.md.

Notes

A different rotation speed for each pair

In /root/work/tf-rope/rope.py, create DIM = 64, THETA_BASE = 10000.0, demo_vectors(), dot(a, b) and thetas(d, base=THETA_BASE). thetas is a list of length d // 2 and its i-th value is base ** (-2 * i / d). demo_vectors() returns two lists of length DIM with q[j] = math.cos(0.7 * j + 0.3) and k[j] = math.sin(0.4 * j + 1.1).

There is one per pair, not per dimension — with 64 dimensions, that is 32. When i is 0 the exponent is 0, so the first value is 1.0, and it shrinks geometrically toward the back. If you do the division with integers, every exponent becomes 0 and all the values are 1.0, so use floating point, like -2.0 * i / d. dot is one line: pair up with zip, multiply, and add.

Rotate once on the plane

Add rotate_pair(x0, x1, angle). It returns a tuple of two, the point (x0, x1) rotated counterclockwise by angle radians. rotate_pair(1.0, 0.0, math.pi / 2) is (0.0, 1.0).

It is (x0*cos - x1*sin, x0*sin + x1*cos). Flipping just one of the two signs makes it clockwise, and then all the values in the later steps change. If the angle is 0 it must give back the original point, and for any angle the distance from the origin must not change — because the squares of cos and sin add up to 1.

Rotate a vector by its position

Add apply_rope(vec, pos, base=THETA_BASE). Group pairs with neighbors, as in (vec[0], vec[1]), (vec[2], vec[3]), and return a new list in which the i-th pair is rotated by pos * thetas(len(vec), base)[i]. Leave the list you received as it is.

Each pair has a different angle — if you use the angle of pair 0 for all pairs, it is just rotating the whole thing, and only a single layer of position information goes in. If pos is 0, all the angles are 0 and the result must equal the original vector, and for any position the norm of the vector must not change. Do not forget to take base and pass it to thetas — if you use only the default, it is silently wrong when called with a different base.

The score of a rotated query and key

Create rope_score(q, k, m, n, base=THETA_BASE). It returns the floating-point number obtained by rotating the query at position m and the key at position n each and then taking the dot product. It does not divide by √d.

It is one line: dot(apply_rope(q, m, base), apply_rope(k, n, base)). If you rotate only one side, absolute position remains as it is and the property in the later steps collapses. When m and n are equal, you must get the same value as the unrotated dot(q, k) — because both were rotated together in the same direction, so the angle between them stays the same.

Put the additive method alongside

Create sin_pos(pos, d, base=THETA_BASE) and add_score(q, k, m, n, base=THETA_BASE). sin_pos is a position vector of length d, and with angle = pos / (base ** ((2 * (j // 2)) / d)), its j-th value is math.sin(angle) if j is even and math.cos(angle) if j is odd. add_score takes the dot product after adding a position vector to each of q and k.

This side is the method of building values and adding them. If you add to only one side, the comparison does not hold. When you expand the dot product, besides q·k you get q·P[n] and P[m]·k, and since these two terms contain only one of the positions, they cannot be grouped into a difference — you will see the result in numbers in the next step.

The same gap gives the same score

Create offset_scan(q, k, offset, starts). For each position s in starts, put the query at s + offset and the key at s and measure the score with both methods. The return value is a list of (s, 회전 점수, 더하는 점수) tuples (the placeholders are the rotation score and the additive score), in the same order as starts. After building it, print offset_scan(q, k, 2, [3, 10, 100, 4000]) yourself and confirm with your own eyes that the four rotation values are the same while the additive side wobbles.

If you call the rope_score and add_score you built earlier as they are, it is five lines. If you swap the query and key, the sign of the gap flips and you get a different value. The four rotation values should agree to twelve decimal places — they are not exactly the same but a very small difference remains, and that is floating-point error. Compare the scatter on the additive side against that magnitude.

Which pairs remain as it gets farther

Create turns(d, delta, base=THETA_BASE) and slow_pairs(d, delta, base=THETA_BASE). turns is a list of length d // 2, and its i-th value is delta * theta_i / (2 * math.pi), that is, the number of turns that pair makes. slow_pairs is the number of pairs (an integer) whose value is less than 1.0.

One turn is 2π radians. If you forget to divide, you get the angle instead of the number of turns and the counts come out completely different. A pair that has gone past one turn writes that distance and the distance with one turn subtracted as the same angle, so it cannot tell the two apart. Call it with gaps 1, 16, 256 and 4096 and see how the number of remaining pairs decreases.

Leave what you measured as a record

Measure with the q and k of demo_vectors(), write dim, base, offset, starts, rope_scores, rope_spread, add_scores, add_spread, norm_before, norm_max_gap, deltas, slow_pairs, turns_first and turns_last in /root/work/tf-rope/rope_report.json, and write /root/work/tf-rope/rope_report.md in the four sections ## 무엇을 쟀나 ## 같은 간격이면 같은 점수다 ## 더하는 방식은 왜 다른가 ## 멀어지면 어느 성분이 남는가 (the Korean headings mean "What was measured", "The same gap gives the same score", "Why the additive method differs" and "Which component remains as it gets farther"). The gap is 2, the starting positions are 3, 10, 100 and 4000, the positions at which the norm is measured are 0, 1, 7, 100 and 4096, and the gaps at which the turns are counted are 1, 16, 256 and 4096.

Do not write the numbers by hand; fill them in with values obtained by actually running your own code. Take rope_scores and add_scores from the tuples offset_scan returned. rope_spread and add_spread are each the maximum minus the minimum, and norm_max_gap is the largest absolute value of the differences measured at the five positions. turns_first is the number of turns of pair 0 at each gap, and turns_last is the number of turns of the last pair. In the body of the report, write in numbers the first rotation score, the scatter on the additive side and the norm of q — the grader checks whether those three values are in the text.