Transformers — Compute Attention By Hand
An Id Becomes a Vector, a Vector Becomes Scores
In one line
An embedding is just a table you pull rows out of, and the logits on the output side are a multiplication that uses that table once more. The number of values held between the two is the vocabulary size times the width, and the cost of raising the vocabulary is paid there.
Why this was needed
Once tokenization is done, the text becomes a list of integers. But integers cannot be used as they are. It does not mean that token 3 and token 4 are neighbors, but if you leave them as numbers, every operation reads them that way. So a vector is attached to each number. The table that stacks those vectors on one sheet is the embedding table.
The first thing that catches here is why a lookup. Textbooks write "multiply a one-hot vector by a matrix", but the code is the single line table[token_id]. The two seem like different stories. They are not different — in a one-hot, only one position is 1 and the rest are 0, so when you multiply and add, only that one row survives. The values do not differ by a hair; only the number of multiplications differs. A lookup is not a different operation but a shortcut of the same operation.
The second thing that catches is the way out. After passing through all of the attention and blocks, one vector remains. You have to turn it back into words, which is the job of scoring the entire vocabulary. To widen a vector of width d into scores of vocabulary size V, you need a matrix of V×d — and a matrix of that shape already exists. The embedding table on the input side is exactly that shape.
A lookup is only a shortcut of multiplication
The shape of the table is (vocabulary size V, width d). The number of values it holds is V times d. If you keep the width as it is and double only the vocabulary, that count doubles too. You raised the vocabulary to reduce the number of tokens, but the cost comes out of this table.
ids = [7, 7, 41]
rows = [table[i] for i in ids] # 조회
# 같은 값을 원-핫으로 계산하면
one = [0.0] * V; one[7] = 1.0
row = [sum(one[r] * table[r][c] for r in range(V)) for c in range(d)]
# rows[0] 과 row 는 같은 값이다. 곱셈만 V 곱하기 d 번 더 했다.
One more thing shows up here. The first two positions of ids are the same number, so exactly the same vector comes out. It is the same whatever comes before and whatever comes after. An embedding has no context. The work of reading the same word differently depending on position is done later by attention, and the table is just a table. This is also why the first argument of torch.nn.MultiheadAttention is embed_dim — the width of the embedding is the width of the model, so every later layer takes the d fixed in the table as it is.
Where you actually use matrix multiplication, something like NumPy's matmul computes it for you. However, the system Python of this lab Pod has no numpy and it exists only inside /opt/onnx-lab/bin/python, so here you compute the two methods directly with the standard library and check that the values are the same.
The way out: use the same table once more
In the short section dealing with embeddings, Attention Is All You Need notes two things. One is that the two embedding layers and the linear transformation before the softmax share the same weight matrix, and the other is that in the embedding layers, those weights are multiplied by √d. The former is weight tying.
Two things arise from tying. First, the table is one set, so the number of values is half. If kept separate, it is two sets of V·d, which is 2·V·d. Second, the way of scoring becomes a dot product. For a hidden vector h, the score of token t is the dot product of row t of the table with h. So if h becomes equal to the embedding of some token, that token's score becomes the largest — because the dot product with itself is the square of its length, which easily exceeds any other dot product.
This property is convenient, but the pitfall comes from the same place. When tied, one table takes on two jobs at once: the meaning coming in and the score going out. There is no guarantee that an arrangement good for one side is good for the other. So whether to tie or not is not a free choice but a trade — in exchange for halving the parameters, you make the table do two jobs.
Do not find nearby tokens with the dot product
When looking for "tokens close to this token", you must not use the dot product as it is. That is because the dot product has the length of the other vector multiplied into it. Even if the direction matches a little less well, if it is long enough it comes to the front.
Cosine removes that length by dividing it out. It leaves only the direction. The surest way to see the difference is to stretch just one row of the table by some factor. The cosine of that row does not change at all (the direction is the same), and the dot products all grow by that factor. If you pull the neighbor list with cosine, the order stays the same, but if you pull it with the dot product, the stretched row jumps to the front.
So the order of the logits and the "order of closeness in meaning" are not the same thing. The logits are in dot-product order, and length is mixed into them.
Where the paper multiplies by √d
Another line in the same section says the embeddings are multiplied by √d. What changes when you multiply? The direction does not change at all. You multiplied every cell by the same number, so the cosine stays the same. What changes is only the magnitude, and the magnitude becomes exactly √d times.
Why does magnitude matter? Where positional information is added to the embedding, if the magnitudes of the two signals are too different, one side gets buried. That is why it is necessary to match the magnitudes. In this lab, instead of proving "why exactly √d", you confirm in numbers that multiplying makes the magnitude √d times and leaves the direction as it is. That much is what you can say by measuring.
What it looks like in the field
First, a proposal to raise the vocabulary ends in a memory meeting. You think it would be good because tokens would decrease and propose doubling the vocabulary, and the answer comes back that the table doubles. It is so even though you did not touch the width. Which side is the gain has to be found by counting the two values side by side.
Second, the parameter counts disagree over "tied or not tied". Even with the same configuration written down, the computed values differ by one table. It is a whole table being there or not, so it is nothing like a rounding error.
Third, the similar-token list is strange. If you pull it with the dot product and call it "close in meaning", the rows with a large length intrude on every query. If you switch to cosine, those rows disappear.
Fourth, you take out only the embedding and expect context. The same word is always the same row in the table. If you need a vector that represents a sentence, you have to pass it through the model, and the value from looking up the table is a value with no context.
Fifth, if you change the width, everything after it moves along. The d of the embedding is the width that every later layer takes, so you cannot fix just one place.
What really matters in practice
- Count the shape of the table first. Multiplying the two numbers V and d gives the space that table occupies. A discussion about raising the vocabulary should begin by writing down this product.
- Check first whether it is tied. It is the first place to look when the parameter counts disagree.
- Closeness by cosine, scores by dot product. If you mix the two, length silently changes the order.
- Know that a lookup and a multiplication are the same operation. You stop being afraid of the optimization, and conversely you do not give the wrong explanation that "it differs because it is a lookup".
What you will do in the next lab
You grow /root/work/tf-embed/embed.py one step at a time. You use only the standard library — the system Python of this Pod has no numpy, torch or transformers, and numpy exists only inside /opt/onnx-lab/bin/python. You do not call a real model, so you do not use numbers such as the vocabulary size or parameter count of a real model. Every value that comes out here was measured from the table you built.
You start by building a deterministic embedding table, count its shape and parameter count, pull out rows by number, and recompute the same values with a one-hot multiplication to see whether the two values are the same. Then you produce logits with the same table, count the number of values when tied and when kept separate, and try doubling the vocabulary.
The last two steps are the point. You stretch just one row of the table by some factor and pull the neighbor list with cosine and with the dot product. The cosine list stays the same, while in the dot-product list the stretched row comes to the front. Finally you multiply the embedding by √d and measure side by side that the magnitude becomes exactly √d times and that the direction stays the same. The grader actually imports your module, pokes at the functions with a different table and different numbers every time, and checks them against values it computes separately.