TT Lab
Get started
Learn Learning paths Courses

Transformers — Compute Attention By Hand

Build Temperature, top-k and top-p by Hand

Continue in TT Lab

Goal

From one set of logits up to choosing the next single token, build every place the distribution passes through using only the standard library. That is a stable softmax after dividing by the temperature, entropy, top-k, top-p, a function that chains the four in a fixed order, inverse cumulative distribution sampling, and greedy and a temperature sweep. At the end, record whether drawing twice with the same seed gives the same sequence and whether the lower-temperature side becomes the same as greedy.

Why it matters

What a transformer block does ends at putting out one list of real numbers the size of the vocabulary. Choosing the token to actually use is a rule outside the model, and both the reason the same model speaks differently each time and the reason it repeats the same thing each time are entirely there. This lab does not call a model. This Pod has neither transformers nor torch, and numpy exists only inside /opt/onnx-lab/bin/python, so import numpy does not work in the system Python. Instead you write down one set of logits yourself and build everything after that by hand. So every number that appears here was measured from the logits you wrote. The difficulty is in the details. Whether the temperature divides before or after the softmax, why the maximum is subtracted, how top-k ties are broken, whether top-p includes the entry that reaches p, and in which order you apply temperature, top-p and top-k. If even one differs, the same settings become a different system. The grader does not trust the explanations you wrote down. It actually imports your module, pokes at the functions with different logits and different temperatures every time, and checks them against probability vectors it computes separately. It does not judge random sampling itself — it looks only at the sequence of indices with the seed fixed and the probability vectors.

Steps

  1. In /root/work/tf-sample/sample.py, create VOCAB, LOGITS and softmax_t(logits, temperature). Divide the logits by the temperature, then subtract the maximum and apply the softmax.
  2. Add entropy_bits(probs) so that it measures how spread out the distribution is as entropy with base 2. Positions with probability 0 count as 0.
  3. Add top_k_filter(probs, k) so that it keeps only the k with the largest probabilities, sets the rest to exactly 0.0, and normalizes again. In case of a tie, the smaller number stays.
  4. Add top_p_filter(probs, p) so that it keeps up to the first entry at which the descending cumulative sum reaches p. If you leave out that entry, the sum falls short of p.
  5. Create filtered_probs(logits, temperature, top_k=None, top_p=None) so that it chains them in the order temperature → softmax → top-k → top-p. If a parameter is None, that step is skipped.
  6. Create sample_index(probs, u) and sample_sequence(logits, temperature, top_k, top_p, n, seed). Draw by inverse cumulative distribution, and call random() of random.Random(seed) n times so that the same seed gives the same sequence.
  7. Create greedy(logits) and sweep(logits, temps, p). greedy is the index of the largest logit (the smaller index in case of a tie), and sweep measures (온도, 엔트로피, top-p 가 남기는 개수) for each temperature (the placeholders are the temperature, the entropy and the number top-p keeps).
  8. Measure with the fixed settings and record the results in /root/work/tf-sample/sample_report.json and /root/work/tf-sample/sample_report.md.

Notes

Softmax after dividing by the temperature

In /root/work/tf-sample/sample.py, set VOCAB (at least 12 distinct short strings) and LOGITS (the same length, all values different, a gap of at least 0.5 between 1st and 2nd place, and a gap of at least 3.0 between the maximum and the minimum), and create softmax_t(logits, temperature). Divide the logits by temperature, then subtract the maximum, take the exponential and normalize. If temperature is 0 or below, raise an exception.

Where you divide is the point. If you compute the probabilities first and then apply the temperature to them, you get a completely different result. Subtracting the maximum comes alive when the temperature is small — dividing by 0.01 turns a logit of 5.2 into 520, and math.exp(520) overflows outright. You are subtracting the same number from the same values, so the probabilities do not change. Temperature 0 cannot be divided by, so it is better to raise a ValueError — greedy is not 'temperature 0' but a separate rule.

Measure the spread in bits

Add entropy_bits(probs). It is entropy with base 2, -sum(p * log2(p)). Positions with probability 0 count as 0. For n entries spread evenly, you get log2(n).

math.log2(0) raises an exception. That is why you must skip positions that are 0, and this is not a hack but the correct handling — because the limit of p * log2(p) as p goes to 0 is 0. If you measure this value while raising the temperature, you can see it grow. This is where you summarize how spread out the distribution is in a single number.

Keep only the top k

Add top_k_filter(probs, k). Keep only the k with the largest probabilities, set the rest to exactly 0.0, and normalize what is left so that it sums to 1. In case of a tie, the smaller number stays. If k is at least the length of the list, discard nothing and only normalize, and if k is less than 1, raise an exception.

If you sort the indices by probability in descending order and hold the first k as a set, the rest is just checking against that set. If you make the sort key (-확률, 번호) (the placeholders are the probability and the index), the tie rule fits in the same line. If you forget to normalize, the sum is not 1, and then when you draw later it gets pulled toward the last position. You must not set discarded positions to a very small value — because they come back to life with a low probability.

Up to the entry where the cumulative sum reaches p

Add top_p_filter(probs, p). Lay out the probabilities in descending order (the smaller number first in case of a tie), compute the cumulative sum, and stop after including the first entry at which the cumulative sum reaches p or more. The rest are exactly 0.0 and what is left is normalized. If p is 0 or below or greater than 1, raise an exception.

Whether you include it or leave it out is the whole of this step. If you drop the entry that pushed past p, the sum of what remains falls short of p — and then the statement 'keep p worth of mass' does not hold at all. You just stop after putting that index in at the moment the cumulative sum reaches p. However small p is, at least one is left.

Chain the four in a fixed order

Create filtered_probs(logits, temperature, top_k=None, top_p=None). Chain them in the order temperature → softmax → top-k → top-p, and skip a step that is None.

The order changes the result. If you apply temperature first, the distribution itself changes, so even with the same p the number left differs, and if you do top-k first, the survivors are normalized again and the probabilities inflate, so top-p stops earlier. The function is done in four lines — you just call the three functions you built earlier in this order. Do not forget the condition that checks whether top_k or top_p is None.

The same seed gives the same sequence

Create sample_index(probs, u) and sample_sequence(logits, temperature, top_k, top_p, n, seed). sample_index adds up the probabilities in index order and returns the first position at which u is exceeded, and sample_sequence builds the distribution only once, then calls random() of random.Random(seed) n times and returns n indices.

It is the first position where u < 누적합 becomes true (the placeholder is the cumulative sum). A position with probability 0.0 does not increase the cumulative sum, so it is never drawn — that means a token that was cut off does not come back to life, and this is the safeguard of this structure. Create random.Random(seed) only once inside the function. If you create a new generator every time you draw, you get the same index n times. Build the distribution only once, too, outside the loop.

Greedy and a temperature sweep

Create greedy(logits) and sweep(logits, temps, p). greedy is the index of the largest logit, and the smaller index in case of a tie. sweep returns a list of three-element tuples (온도, 엔트로피, top-p 가 남기는 개수) for each temperature in temps (the placeholders are the temperature, the entropy and the number top-p keeps).

greedy does not even need to compute the probabilities — because the softmax does not change the order. If you are unsure which side max gives in a tie, sweep through by hand and update only when strictly larger. Then the smaller index stays. The count in sweep is the number of nonzero positions after passing through top_p_filter. See with your own eyes that as the temperature rises, the entropy grows and the number kept grows too.

Leave the result of turning the knobs

Measure with the fixed settings. The seed is 20260917, the number of draws is 24, top_k is 5 and top_p is 0.9, and the hot side uses temperature 1.0 and the cold side temperature 0.2. Write the results in /root/work/tf-sample/sample_report.json as vocab_size, greedy_index, greedy_token, top_prob_t1, entropy_t1, sweep, nucleus, topk_mass, seed, draw_count, hot_draws, cold_draws, cold_is_greedy, hot_distinct, cold_distinct and repeat_matches, and write /root/work/tf-sample/sample_report.md in the four sections ## 무엇을 쟀나 ## 온도가 분포를 어떻게 바꾸나 ## top-k 와 top-p 가 남기는 것 ## 같은 시드는 같은 수열을 준다 (the Korean headings mean "What was measured", "How temperature changes the distribution", "What top-k and top-p leave" and "The same seed gives the same sequence").

Do not write the numbers by hand; fill them in with values obtained by actually running your own code. sweep uses temperatures 0.25, 0.5, 1.0, 2.0 and 4.0 with p of 0.9. nucleus is [p, 남은 개수] measured at temperature 1.0 while changing p through 0.5, 0.8, 0.9 and 0.95 (the placeholders are p and the number kept), and topk_mass is [k, 합], the sum occupied by the top k in the original probabilities at temperature 1.0, measured for k of 1, 3, 5 and 10 (the placeholders are k and the sum) — not the value normalized after cutting but the sum before cutting. cold_is_greedy is whether the cold-side sequence is entirely equal to greedy(LOGITS), and repeat_matches is whether drawing once more with the same seed gives the same sequence. Depending on the logits you wrote, cold_is_greedy may be false — if so, write it as it is and explain why in the text.