Transformers — Compute Attention By Hand
Attention Is a Weighted Average
In one line
Attention is an operation that decides "from this position, how much to look at each position" and takes a weighted average of the values. That is all it is.
Three names
Three vectors are made from the same input.
- Q (query) — what I am looking for
- K (key) — the label each position puts forward
- V (value) — the content that position actually offers
The library analogy is accurate. You take a question (Q), match it against the titles on the book spines (K), and bring back the contents (V) of the books that match well, in proportion to how well they match.
score = Q · Kᵀ / √d ← 얼마나 맞는가
weight = softmax(score) ← 합이 1이 되게
out = weight · V ← 가중 평균
That is three lines. Everything else is just doing these three lines many times, in many directions.
Why divide by √d — in numbers
This is the part people most often skip with "that's just how it's done", but the reason is clear.
The dot product of two d-dimensional vectors with independent components of variance 1 has a variance of d. With d=64, the scores spread over roughly a ±8 range.
What happens when scores of ±8 go into the softmax? e⁸ / e⁻⁸ ≈ 9백만 (the Korean number in the code means 9 million). The largest one takes almost all of the 1, and the rest become 0. It is no longer a weighted average, just "pick one".
Then training does not work. When the softmax saturates, the gradient gets close to 0.
Dividing by √d brings the variance back to 1. In step 2 of this lab you measure this yourself — how high the maximum probability climbs when you do not divide.
Softmax is computed after subtracting the maximum
def softmax(xs):
m = max(xs) # 이 한 줄이 없으면
e = [exp(x - m) for x in xs] # exp(1000) 에서 터진다
s = sum(e)
return [v / s for v in e]
exp(x - max) is mathematically the same value, but it does not overflow. In production code, if you leave out this one line, inf and nan appear the moment a large value comes in.
The mask — so it cannot see the future
A language model is trained to predict the next token. But attention by default looks at every position. That would be predicting the answer while looking at the answer.
So the upper triangle of the score matrix is set to -inf.
k0 k1 k2
q0 0.3 -inf -inf
q1 0.1 0.5 -inf
q2 0.2 0.1 0.4
-inf becomes exactly 0 after passing through exp. The point is to add -inf before the softmax, not to multiply by 0 — if you multiply by 0 after the softmax, the remaining weights no longer sum to 1.
This is the "causal" or "decoder" mask. Encoder models like BERT do not have it. That is why BERT is strong at reading and understanding a whole sentence, and GPT is strong at continuing text.
Multi-head — why split
Instead of using d=64 all at once, you split it into 8 and run attention with d=8 eight times, then concatenate the results. The amount of computation is almost the same.
Why? A single softmax can express only one kind of relationship. Because the weights sum to 1, it cannot look strongly at several places at once. Once you split the heads, one head looks at the word right before, another at the subject at the start of the sentence, another at the matching quotation mark.
If you open up the heads of a trained model, they really are divided up that way.
Positional encoding — attention does not know order
This is the most surprising part when you first learn it.
Attention has no notion of order at all. If you shuffle the input tokens, the output is simply shuffled in the same way, and the values themselves do not change (permutation equivariant). "I like you" and "you I like" cannot be told apart.
So positional information is added to the input.
- Sine/cosine (the original paper) — has no learned parameters and extrapolates to longer lengths
- Learned embeddings (BERT, ViT) — simple, but cannot go beyond the length seen in training
- RoPE (LLaMA, Qwen and most models today) — rotates Q·K so that relative position is captured naturally. It extends to longer lengths easily, so the models of today that stretch their context all use it
- ALiBi — adds a penalty proportional to distance to the score
In step 5 you check this yourself. If you shuffle the input without positional encoding, the output is shuffled right along with it; if you add it, that no longer holds.
LayerNorm and residual connections
x = x + Attention(LayerNorm(x))
x = x + FFN(LayerNorm(x))
- Residual — the gradient reaches the input even when you stack many layers. Without it, training fails once you go past about 6 layers
- LayerNorm — sets each token vector to mean 0 and variance 1. The difference from BatchNorm is that it normalizes along the feature axis, not the batch, so it is not affected by batch size or sequence length
Whether you put LayerNorm before or after attention (pre-LN vs post-LN) actually decides training stability. The original paper was post-LN, but nowadays nearly everything is pre-LN — because it trains without warmup.
Where the cost comes from
The score matrix is n × n. Memory and computation grow with the square of the sequence length. Raising the context from 4k to 8k makes it 4 times as large.
Attempts to reduce this are a major line of research today.
- FlashAttention — keeps the math as it is and changes the order of GPU memory access, raising real speed several times over. It is not an approximation
- GQA / MQA — several Q heads share K·V heads, which reduces the KV cache at inference time. Most open models today use GQA
- MoE — turns on only some of several experts in each layer, keeping the parameters large but the computation small
- Sliding window / sparse attention — does not look at distant positions at all
Summary
Attention itself is three lines. The rest is about how to do those three lines stably, cheaply, and while knowing the order. In the next lab you write those three lines yourself and see in numbers what breaks without scaling and positional encoding.
In the field
You will rarely have to write this computation yourself. The framework does it all. The reason you still need to know it is that where to look when something goes wrong is decided here.
When the training loss suddenly becomes nan, it is usually because the values before the softmax overflowed. If the mask is wrong, training quietly looks at future tokens without any error, and only the evaluation score comes out strangely good. If the results change when you change the batch size, suspect the normalization axis. And if you double the context and memory jumps fourfold, that is not a bug; it is natural given the structure.
If you run an inference service, the KV cache is your memory budget. First calculate whether the number of concurrent requests multiplied by the maximum context length fits in GPU memory, and if it does not, the real options are to choose a model that uses GQA or to lower the context limit.