Transformers — Compute Attention By Hand
Build Temperature, top-k and top-p by Hand
Goal
From one set of logits up to choosing the next single token, build every place the distribution passes through using only the standard library. That is a stable softmax after dividing by the temperature, entropy, top-k, top-p, a function that chains the four in a fixed order, inverse cumulative distribution sampling, and greedy and a temperature sweep. At the end, record whether drawing twice with the same seed gives the same sequence and whether the lower-temperature side becomes the same as greedy.
Why it matters
What a transformer block does ends at putting out one list of real numbers the size of the vocabulary. Choosing the token to actually use is a rule outside the model, and both the reason the same model speaks differently each time and the reason it repeats the same thing each time are entirely there.
This lab does not call a model. This Pod has neither transformers nor torch, and numpy exists only inside /opt/onnx-lab/bin/python, so import numpy does not work in the system Python. Instead you write down one set of logits yourself and build everything after that by hand. So every number that appears here was measured from the logits you wrote.
The difficulty is in the details. Whether the temperature divides before or after the softmax, why the maximum is subtracted, how top-k ties are broken, whether top-p includes the entry that reaches p, and in which order you apply temperature, top-p and top-k. If even one differs, the same settings become a different system.
The grader does not trust the explanations you wrote down. It actually imports your module, pokes at the functions with different logits and different temperatures every time, and checks them against probability vectors it computes separately. It does not judge random sampling itself — it looks only at the sequence of indices with the seed fixed and the probability vectors.
Steps
- In /root/work/tf-sample/sample.py, create
VOCAB,LOGITSandsoftmax_t(logits, temperature). Divide the logits by the temperature, then subtract the maximum and apply the softmax. - Add
entropy_bits(probs)so that it measures how spread out the distribution is as entropy with base 2. Positions with probability 0 count as 0. - Add
top_k_filter(probs, k)so that it keeps only the k with the largest probabilities, sets the rest to exactly0.0, and normalizes again. In case of a tie, the smaller number stays. - Add
top_p_filter(probs, p)so that it keeps up to the first entry at which the descending cumulative sum reaches p. If you leave out that entry, the sum falls short of p. - Create
filtered_probs(logits, temperature, top_k=None, top_p=None)so that it chains them in the order temperature → softmax → top-k → top-p. If a parameter isNone, that step is skipped. - Create
sample_index(probs, u)andsample_sequence(logits, temperature, top_k, top_p, n, seed). Draw by inverse cumulative distribution, and callrandom()ofrandom.Random(seed)n times so that the same seed gives the same sequence. - Create
greedy(logits)andsweep(logits, temps, p).greedyis the index of the largest logit (the smaller index in case of a tie), andsweepmeasures(온도, 엔트로피, top-p 가 남기는 개수)for each temperature (the placeholders are the temperature, the entropy and the number top-p keeps). - Measure with the fixed settings and record the results in /root/work/tf-sample/sample_report.json and /root/work/tf-sample/sample_report.md.
Notes
- Execution contract: the grader imports
/root/work/tf-sample/sample.pyas a Python module and usesVOCAB,LOGITS,softmax_t,entropy_bits,top_k_filter,top_p_filter,filtered_probs,sample_index,sample_sequence,greedyandsweepdirectly. It does not run it as a script, soif __name__ == "__main__"is not needed. VOCABis a list of at least 12 distinct short strings. Choose the content freely.LOGITSis a list of real numbers of the same length asVOCAB, and the values must all differ. They are unnormalized scores, not probabilities. The gap between 1st and 2nd place must be at least 0.5, and the gap between the maximum and the minimum must be at least 3.0, so that you can see the distribution move when you change the temperature.softmax_t(logits, temperature)first divides the logits bytemperature, subtracts the maximum from the divided values, takes the exponential, and normalizes. If you do not subtract the maximum,math.expoverflows when the temperature is small. Iftemperatureis 0 or below, raise an exception.entropy_bits(probs)is-sum(p * log2(p)). Skip positions wherepis 0. For n entries spread evenly, you getlog2(n).top_k_filter(probs, k)sets discarded positions to exactly0.0and normalizes only what is left so that it sums to 1. In case of a tie, the smaller number stays. Ifkis at least the length of the list, it discards nothing and only normalizes. Ifkis less than 1, raise an exception.top_p_filter(probs, p)lays out the probabilities in descending order (the smaller number first in case of a tie), computes the cumulative sum, and stops after including the first entry at which the cumulative sum reaches p or more. The rest are exactly0.0and what is left is normalized.pmust be greater than 0 and at most 1, and otherwise raise an exception.- The order of
filtered_probsis temperature → softmax → top-k → top-p. This order is the contract of this lab, and if you change the order, the tokens that remain change. sample_index(probs, u)adds up the probabilities in index order and returns the first position at whichuis exceeded (the first position whereu < 누적합becomes true, where the placeholder is the cumulative sum).uis at least 0 and less than 1. It never returns a position whose probability is0.0.sample_sequence(logits, temperature, top_k, top_p, n, seed)builds the distribution only once, callsrandom()ofrandom.Random(seed)n times, and returns a list of n indices. The same seed must give the same result.sweep(logits, temps, p)returns a list of three-element tuples(온도, 엔트로피, top-p 가 남기는 개수)for each temperature intemps(the placeholders are the temperature, the entropy and the number top-p keeps). The count is the number of nonzero positions after passing throughtop_p_filter.- The step 8 settings are fixed. The seed is 20260917, the number of draws is 24,
top_kis 5 andtop_pis 0.9. The hot side draws at temperature 1.0 and the cold side at temperature 0.2, from the same seed. - The
sweeptemperatures in step 8 are 0.25 · 0.5 · 1.0 · 2.0 · 4.0 andpis 0.9.nucleusis[p, 남은 개수]measured at temperature 1.0 while changing p through 0.5 · 0.8 · 0.9 · 0.95 (the placeholders are p and the number kept), andtopk_massis[k, 합], the sum occupied by the top k in the original probabilities at temperature 1.0, measured for k of 1 · 3 · 5 · 10 (the placeholders are k and the sum). - The keys of
sample_report.json:vocab_size,greedy_index,greedy_token,top_prob_t1,entropy_t1,sweep,nucleus,topk_mass,seed,draw_count,hot_draws,cold_draws,cold_is_greedy,hot_distinct,cold_distinct,repeat_matches. sample_report.mdis written in the four sections## 무엇을 쟀나## 온도가 분포를 어떻게 바꾸나## top-k 와 top-p 가 남기는 것## 같은 시드는 같은 수열을 준다(the Korean headings mean "What was measured", "How temperature changes the distribution", "What top-k and top-p leave" and "The same seed gives the same sequence").- This Pod has no internet.
pip installdoes not work and there is neither transformers nor torch. numpy exists only inside/opt/onnx-lab/bin/python, soimport numpydoes not work in the system Python.mathandrandomare enough. - Official documents: the original nucleus sampling paper · the original top-k paper · Hugging Face — Generation strategies · Python — math
- Common mistakes: applying the temperature to the probabilities, skipping the subtraction of the maximum and overflowing at low temperature, forgetting to normalize after cutting, dropping the first entry that exceeds p in top-p, not fixing a tie rule, applying top-p before top-k, and rebuilding the distribution every time you draw.
Softmax after dividing by the temperature
In /root/work/tf-sample/sample.py, set VOCAB (at least 12 distinct short strings) and LOGITS (the same length, all values different, a gap of at least 0.5 between 1st and 2nd place, and a gap of at least 3.0 between the maximum and the minimum), and create softmax_t(logits, temperature). Divide the logits by temperature, then subtract the maximum, take the exponential and normalize. If temperature is 0 or below, raise an exception.
Where you divide is the point. If you compute the probabilities first and then apply the temperature to them, you get a completely different result. Subtracting the maximum comes alive when the temperature is small — dividing by 0.01 turns a logit of 5.2 into 520, and math.exp(520) overflows outright. You are subtracting the same number from the same values, so the probabilities do not change. Temperature 0 cannot be divided by, so it is better to raise a ValueError — greedy is not 'temperature 0' but a separate rule.
Measure the spread in bits
Add entropy_bits(probs). It is entropy with base 2, -sum(p * log2(p)). Positions with probability 0 count as 0. For n entries spread evenly, you get log2(n).
math.log2(0) raises an exception. That is why you must skip positions that are 0, and this is not a hack but the correct handling — because the limit of p * log2(p) as p goes to 0 is 0. If you measure this value while raising the temperature, you can see it grow. This is where you summarize how spread out the distribution is in a single number.
Keep only the top k
Add top_k_filter(probs, k). Keep only the k with the largest probabilities, set the rest to exactly 0.0, and normalize what is left so that it sums to 1. In case of a tie, the smaller number stays. If k is at least the length of the list, discard nothing and only normalize, and if k is less than 1, raise an exception.
If you sort the indices by probability in descending order and hold the first k as a set, the rest is just checking against that set. If you make the sort key (-확률, 번호) (the placeholders are the probability and the index), the tie rule fits in the same line. If you forget to normalize, the sum is not 1, and then when you draw later it gets pulled toward the last position. You must not set discarded positions to a very small value — because they come back to life with a low probability.
Up to the entry where the cumulative sum reaches p
Add top_p_filter(probs, p). Lay out the probabilities in descending order (the smaller number first in case of a tie), compute the cumulative sum, and stop after including the first entry at which the cumulative sum reaches p or more. The rest are exactly 0.0 and what is left is normalized. If p is 0 or below or greater than 1, raise an exception.
Whether you include it or leave it out is the whole of this step. If you drop the entry that pushed past p, the sum of what remains falls short of p — and then the statement 'keep p worth of mass' does not hold at all. You just stop after putting that index in at the moment the cumulative sum reaches p. However small p is, at least one is left.
Chain the four in a fixed order
Create filtered_probs(logits, temperature, top_k=None, top_p=None). Chain them in the order temperature → softmax → top-k → top-p, and skip a step that is None.
The order changes the result. If you apply temperature first, the distribution itself changes, so even with the same p the number left differs, and if you do top-k first, the survivors are normalized again and the probabilities inflate, so top-p stops earlier. The function is done in four lines — you just call the three functions you built earlier in this order. Do not forget the condition that checks whether top_k or top_p is None.
The same seed gives the same sequence
Create sample_index(probs, u) and sample_sequence(logits, temperature, top_k, top_p, n, seed). sample_index adds up the probabilities in index order and returns the first position at which u is exceeded, and sample_sequence builds the distribution only once, then calls random() of random.Random(seed) n times and returns n indices.
It is the first position where u < 누적합 becomes true (the placeholder is the cumulative sum). A position with probability 0.0 does not increase the cumulative sum, so it is never drawn — that means a token that was cut off does not come back to life, and this is the safeguard of this structure. Create random.Random(seed) only once inside the function. If you create a new generator every time you draw, you get the same index n times. Build the distribution only once, too, outside the loop.
Greedy and a temperature sweep
Create greedy(logits) and sweep(logits, temps, p). greedy is the index of the largest logit, and the smaller index in case of a tie. sweep returns a list of three-element tuples (온도, 엔트로피, top-p 가 남기는 개수) for each temperature in temps (the placeholders are the temperature, the entropy and the number top-p keeps).
greedy does not even need to compute the probabilities — because the softmax does not change the order. If you are unsure which side max gives in a tie, sweep through by hand and update only when strictly larger. Then the smaller index stays. The count in sweep is the number of nonzero positions after passing through top_p_filter. See with your own eyes that as the temperature rises, the entropy grows and the number kept grows too.
Leave the result of turning the knobs
Measure with the fixed settings. The seed is 20260917, the number of draws is 24, top_k is 5 and top_p is 0.9, and the hot side uses temperature 1.0 and the cold side temperature 0.2. Write the results in /root/work/tf-sample/sample_report.json as vocab_size, greedy_index, greedy_token, top_prob_t1, entropy_t1, sweep, nucleus, topk_mass, seed, draw_count, hot_draws, cold_draws, cold_is_greedy, hot_distinct, cold_distinct and repeat_matches, and write /root/work/tf-sample/sample_report.md in the four sections ## 무엇을 쟀나 ## 온도가 분포를 어떻게 바꾸나 ## top-k 와 top-p 가 남기는 것 ## 같은 시드는 같은 수열을 준다 (the Korean headings mean "What was measured", "How temperature changes the distribution", "What top-k and top-p leave" and "The same seed gives the same sequence").
Do not write the numbers by hand; fill them in with values obtained by actually running your own code. sweep uses temperatures 0.25, 0.5, 1.0, 2.0 and 4.0 with p of 0.9. nucleus is [p, 남은 개수] measured at temperature 1.0 while changing p through 0.5, 0.8, 0.9 and 0.95 (the placeholders are p and the number kept), and topk_mass is [k, 합], the sum occupied by the top k in the original probabilities at temperature 1.0, measured for k of 1, 3, 5 and 10 (the placeholders are k and the sum) — not the value normalized after cutting but the sum before cutting. cold_is_greedy is whether the cold-side sequence is entirely equal to greedy(LOGITS), and repeat_matches is whether drawing once more with the same seed gives the same sequence. Depending on the logits you wrote, cold_is_greedy may be false — if so, write it as it is and explain why in the text.