TT Lab
Get started
Learn Learning paths Courses

Transformers — Compute Attention By Hand

Confirm with numbers why you divide by √d

고급 · Lessons 31 · Lab 10

Start the lab

Curriculum

Inside Attention

Tokenization and Vocabulary

Rotary Position Embedding (RoPE)

Temperature, top-k and top-p

Context Length and Cost

What a KV Cache Reduces

Embeddings, Weight Tying, and Logits

How Integer Quantization Hits Accuracy

MHA, MQA, and GQA

Cross Entropy and Perplexity

Reference docs