TT Lab
Get started
Learn Learning paths Courses

Transformers — Compute Attention By Hand

Attention Is a Weighted Average

Continue in TT Lab

In one line

Attention is an operation that decides "from this position, how much to look at each position" and takes a weighted average of the values. That is all it is.

Three names

Three vectors are made from the same input.

The library analogy is accurate. You take a question (Q), match it against the titles on the book spines (K), and bring back the contents (V) of the books that match well, in proportion to how well they match.

score  = Q · Kᵀ / √d       ← 얼마나 맞는가
weight = softmax(score)     ← 합이 1이 되게
out    = weight · V         ← 가중 평균

That is three lines. Everything else is just doing these three lines many times, in many directions.

Why divide by √d — in numbers

This is the part people most often skip with "that's just how it's done", but the reason is clear.

The dot product of two d-dimensional vectors with independent components of variance 1 has a variance of d. With d=64, the scores spread over roughly a ±8 range.

What happens when scores of ±8 go into the softmax? e⁸ / e⁻⁸ ≈ 9백만 (the Korean number in the code means 9 million). The largest one takes almost all of the 1, and the rest become 0. It is no longer a weighted average, just "pick one".

Then training does not work. When the softmax saturates, the gradient gets close to 0.

Dividing by √d brings the variance back to 1. In step 2 of this lab you measure this yourself — how high the maximum probability climbs when you do not divide.

Softmax is computed after subtracting the maximum

def softmax(xs):
    m = max(xs)                       # 이 한 줄이 없으면
    e = [exp(x - m) for x in xs]      # exp(1000) 에서 터진다
    s = sum(e)
    return [v / s for v in e]

exp(x - max) is mathematically the same value, but it does not overflow. In production code, if you leave out this one line, inf and nan appear the moment a large value comes in.

The mask — so it cannot see the future

A language model is trained to predict the next token. But attention by default looks at every position. That would be predicting the answer while looking at the answer.

So the upper triangle of the score matrix is set to -inf.

      k0    k1    k2
q0   0.3  -inf  -inf
q1   0.1   0.5  -inf
q2   0.2   0.1   0.4

-inf becomes exactly 0 after passing through exp. The point is to add -inf before the softmax, not to multiply by 0 — if you multiply by 0 after the softmax, the remaining weights no longer sum to 1.

This is the "causal" or "decoder" mask. Encoder models like BERT do not have it. That is why BERT is strong at reading and understanding a whole sentence, and GPT is strong at continuing text.

Multi-head — why split

Instead of using d=64 all at once, you split it into 8 and run attention with d=8 eight times, then concatenate the results. The amount of computation is almost the same.

Why? A single softmax can express only one kind of relationship. Because the weights sum to 1, it cannot look strongly at several places at once. Once you split the heads, one head looks at the word right before, another at the subject at the start of the sentence, another at the matching quotation mark.

If you open up the heads of a trained model, they really are divided up that way.

Positional encoding — attention does not know order

This is the most surprising part when you first learn it.

Attention has no notion of order at all. If you shuffle the input tokens, the output is simply shuffled in the same way, and the values themselves do not change (permutation equivariant). "I like you" and "you I like" cannot be told apart.

So positional information is added to the input.

In step 5 you check this yourself. If you shuffle the input without positional encoding, the output is shuffled right along with it; if you add it, that no longer holds.

LayerNorm and residual connections

x = x + Attention(LayerNorm(x))
x = x + FFN(LayerNorm(x))

Whether you put LayerNorm before or after attention (pre-LN vs post-LN) actually decides training stability. The original paper was post-LN, but nowadays nearly everything is pre-LN — because it trains without warmup.

Where the cost comes from

The score matrix is n × n. Memory and computation grow with the square of the sequence length. Raising the context from 4k to 8k makes it 4 times as large.

Attempts to reduce this are a major line of research today.

Summary

Attention itself is three lines. The rest is about how to do those three lines stably, cheaply, and while knowing the order. In the next lab you write those three lines yourself and see in numbers what breaks without scaling and positional encoding.

In the field

You will rarely have to write this computation yourself. The framework does it all. The reason you still need to know it is that where to look when something goes wrong is decided here.

When the training loss suddenly becomes nan, it is usually because the values before the softmax overflowed. If the mask is wrong, training quietly looks at future tokens without any error, and only the evaluation score comes out strangely good. If the results change when you change the batch size, suspect the normalization axis. And if you double the context and memory jumps fourfold, that is not a bug; it is natural given the structure.

If you run an inference service, the KV cache is your memory budget. First calculate whether the number of concurrent requests multiplied by the maximum context length fits in GPU memory, and if it does not, the real options are to choose a model that uses GQA or to lower the context limit.