Transformers — Compute Attention By Hand
The Current Model Landscape and the Encoders
In one line
The models in use as of 2026 fall along three axes: which mask they use, what they treat as a token, and how they save computation.
Why split them up — axis 1, three families by mask
| Family | Mask | Good at | Examples |
|---|---|---|---|
| Decoder | Causal (cannot see what comes after) | Continuing text, conversation, code | GPT, LLaMA, Claude, Qwen |
| Encoder | None (bidirectional) | Classification, embeddings for search | BERT, RoBERTa, E5 |
| Encoder-decoder | Both | Translation, summarization, speech recognition | T5, Whisper |
One fork in the road that matters in practice. Use an encoder model for the embeddings in search (RAG). Using the last hidden state of a generative decoder model as an embedding usually performs worse — because of the causal mask, only the later tokens see what came before.
Axis 2 — what counts as a token
A transformer is not text-only. Anything that can be made into a "sequence of vectors" goes in. That is why the encoder is needed.
Vision encoders (the ViT family)
An image is cut into patches and turned into tokens.
- Cutting a 224×224 image into 16×16 patches gives 196 tokens
- Flattening each patch (768 dimensions) and applying a linear transformation gives the embedding
- Positional encoding is added — without it, shuffling the patches gives the same result (the very property from the previous lab)
The difference from a CNN is that there is no locality assumption. A CNN has built in from the start the structure of looking at nearby pixels first. A ViT can see the whole image from the first layer. That is why a CNN is better when data is scarce, and a ViT wins when it is plentiful.
CLIP goes one step further. It has an image encoder and a text encoder separately, and trains them so that matching pairs come close and mismatched pairs move apart. Then images and sentences lie in the same space. This is why you can search for images with the sentence vector for "a photo of a cat", and most multimodal LLMs today attach a CLIP-family vision encoder at the front.
Audio encoders (the Whisper family)
Sound is 16,000 samples per second. It cannot be used as tokens as it is.
- It is converted into a log-mel spectrogram — it becomes a two-dimensional image of time × frequency. The mel scale is a scale that reflects the fact that the human ear is sensitive to low frequencies
- The time axis is reduced with a convolution — Whisper halves it with a conv of stride 2. 30 seconds becomes 1,500 tokens
- A transformer encoder is placed on top
Preprocessing is half of it. How information is compressed at this stage decides performance. It plays exactly the same role as the tokenizer for text.
This is also why Whisper is an encoder-decoder. The encoder understands the whole sound bidirectionally, and the decoder spits out characters one at a time.
So a multimodal LLM is
이미지 → 비전 인코더 → 프로젝션 → [토큰들] ┐
├→ 디코더 LLM → 글자
텍스트 → 토크나이저 → [토큰들] ────────────┘
A single projection layer connects the two worlds. It is a small MLP that matches the output dimension of the vision encoder to the embedding dimension of the LLM. A common approach in training is to align only this part first (the alignment stage) and then unfreeze the whole model little by little.
Axis 3 — how computation is saved
KV cache. There is no need to recompute the K·V of all the preceding tokens each time you generate. You store them and add only the new token's. That is why most of inference memory is the KV cache, not the model weights. The longer the context, the more it dominates.
GQA (grouped-query attention). There are 32 Q heads but only 8 K·V heads, shared four to one. The KV cache becomes a quarter. The quality loss is almost nil, so most open models today use it.
MoE (mixture of experts). Each layer has several FFNs, and only about 2 are turned on per token. The total parameters are large, but the computation one token uses is small. However, everything must be loaded into memory, so this does not mean that serving cost is independent of the parameter count.
Quantization. Weights are reduced to 8 bits or 4 bits. The accuracy loss is smaller than you might think and memory shrinks a lot. However, the KV cache must be quantized separately for it to take effect.
FlashAttention. It is not an approximation. It does not load the large score matrix into memory all at once; it computes in tiles, reducing slow GPU memory accesses. It produces mathematically exactly the same values while being several times faster.
How the context was extended
RoPE is the key. If instead of adding absolute position you rotate Q·K by an angle, the dot product of two positions depends only on the distance. Then it works to some extent even at lengths never seen in training.
- Position interpolation (PI) — compresses the rotation angles to extend a model trained at 4k to 16k
- NTK/YaRN — extends each frequency band differently, hurting short-context performance less
However, a long context is not the same as using it well. The phenomenon of missing information in the middle of a long context (lost in the middle) is well known. Inserting only the pieces you need through retrieval is often more accurate than putting in all 100k.
What you actually look at when choosing
- What is masked — generation or understanding
- How much the KV cache costs — is it GQA, how long is the context
- What the encoder takes in — if you feed in images or sound, the preprocessing of the front-end encoder decides performance
- Whether it breaks when quantized — if the deployment environment is fixed, this is the first question
The parameter count is in none of the four. What actually decides serving cost and quality is the four above.
In the field
The questions that actually come up when choosing a model are not about performance leaderboards but these. Is it all right to pick a decoder model for embeddings used in search? The documents are in Korean, so by how many times does this tokenizer inflate them? It says it supports a 128k context, but what is the actual latency at that length?
And once you know these axes, the speed at which you read vendor documents changes. If it says 'sliding window attention', that means it forgets the earlier part of a long document; 'GQA' means inference memory is small; and 'encoder-decoder' means it is strong at translation and summarization. Instead of learning from scratch every time a new model comes out, you can just place it on these coordinates.
If you are choosing a model to host internally, look at one more thing at the end — the license and how much of the weights are open. However good the performance, if commercial use is blocked, it drops out of the candidates.