TT Lab
Get started
Learn Learning paths Courses

Transformers — Compute Attention By Hand

The Same Prompt, a Different Answer Every Time

Continue in TT Lab

In one line

What the model puts out is not an answer but a set of logits, and the answer is decided in the process of turning those logits into probabilities and drawing one. Temperature, top-k and top-p are three knobs that adjust that probability distribution before drawing.

Why this was needed

When you first attach a model, two odd things happen. One is that you put in the same prompt twice and the answers differ, and the other is that when you draw a long answer it ends up repeating the same thing.

Both come from the last step. What a transformer block does ends at putting out one list of real numbers the size of the vocabulary. This list is called logits, and they are not probabilities but unnormalized scores. Choosing the one token to actually use from here is a rule outside the model, and depending on what that rule is, the same model speaks completely differently.

The simplest rule is to pick the largest logit. It is called greedy. The result is always the same, which makes it good for testing, but when you draw a long output it easily falls into a loop that returns to the same phrase. This is exactly the phenomenon summarized in The Curious Case of Neural Text Degeneration — if you follow only the path with the highest probability, you get text that does not resemble human writing and is noticeably repetitive.

So you draw. If you pick one in proportion to the probabilities, repetition decreases. But now the tail becomes the problem. If the vocabulary has tens of thousands of entries, there are tens of thousands of tokens with a probability of 0.00001, and the sum of that whole tail cannot be ignored. If one of them is drawn even once, the sentence collapses right there. All three knobs are tools for handling this tension — between repetition and collapse.

The point of temperature is where you divide

Temperature T is the value by which you divide the logits before feeding them into the softmax. It is not that you compute the probabilities first and then touch them.

scaled = [value / T for value in logits]
top = max(scaled)                      # 빼는 것도 장식이 아니다
weights = [math.exp(value - top) for value in scaled]
probs = [w / sum(weights) for w in weights]

How the division works is easy to see in terms of gaps. Say the logit gap between 1st and 2nd place is 1.1. If T is 0.25, that gap widens to 4.4, and after taking the exponential, 2nd place's share shrinks to about 1 percent of 1st place's. If T is 4, the gap narrows to 0.275 and the two become almost equal. A small T makes it sharper and a large T makes it flatter. If T is 1, you are using the logits as they are.

The place where you subtract the maximum is not just a convention either. If you set T to 0.01, a logit of 5.2 becomes 520, and math.exp(520) overflows outright. This means the temperature itself produces large inputs. If you subtract the maximum, the argument of the exponential drops to 0 or below so there is nowhere to overflow, and since you are subtracting the same number from the same values, the probabilities do not change.

There is one thing to pin down. There is no softmax with T equal to 0. The division cannot be done. The phrase "set the temperature to 0 and you get greedy" is a shorthand for the limit as T goes to 0, and the actual implementation switches to greedy at that point. A negative value is worse — the ranking is reversed and the least plausible token becomes first.

If you want to see how spread out it is as a single number, use entropy. Entropy with base 2 is close to 0 when the probability is concentrated in one place, and is log2(n) when it is spread evenly over n entries. Raising the temperature makes this value larger.

top-k and top-p are two ways of cutting the tail

top-k is simple. You keep only the k with the largest probabilities, throw away the rest, and normalize what is left again. You have to decide how to break ties. This lab pins it down so that the one with the smaller number stays.

The weakness of fixing k is that it does not look at the shape of the distribution. In a confident position one top entry is enough, yet it keeps k, and at a fork k is not enough.

top-p, also known as nucleus sampling cuts by mass instead of count. You lay out the probabilities in descending order and compute the cumulative sum, and keep up to the first entry at which the cumulative sum reaches p. Here is a spot that is easy to get confused about. If you leave out that entry, the sum of what remains falls short of p. That is why it is "up to the entry that crosses it", not "up to just before crossing". This way the number kept varies with the distribution on its own — one at a sharp position, several at a flat one.

When you use the two together, the order changes the result. If you apply temperature first, the distribution itself changes, so even with the same p the number left differs. If you do top-k first, the survivors are normalized again and the probabilities inflate, and since the cumulative sum is computed with the inflated values, top-p stops earlier. So which order to apply them in is a value you must fix and write in the documentation. Hugging Face's documentation on generation strategies also treats these knobs not separately but as one bundle of settings.

The same seed gives the same sequence

Once you have fixed up the distribution, you draw one. Inverse cumulative distribution is the standard method — you draw one real number u that is at least 0 and less than 1, add up the probabilities in index order, and pick the position where u is first exceeded. A position whose probability is exactly 0 does not increase the cumulative sum, so it is never drawn. That means a token that was cut off does not come back to life.

An important property comes out here. If you fix the seed of the random number generator that gives u, the same input gives the same sequence. It is random yet reproducible. To reproduce an incident, this must be there.

What it looks like in the field

First, you cannot reproduce the report "sometimes a strange answer comes out." It is because the seed was not kept. If you write the seed along with the temperature, k and p in the log for each request, you can bring that one request back exactly.

Second, you try to gain diversity by raising the temperature and lose accuracy. For work with a single answer, such as classification or extraction, there is no reason to raise the temperature. The knobs must be set separately for each job.

Third, the settings are the same but the results differ after switching libraries. The order of application differs, or the tie handling differs, or whether the entry that reaches p is included differs. All three are rarely written down in the documentation.

Fourth, the cache hits at the wrong times. If you use only the prompt as the key and leave out the temperature, k, p and seed, requests with different settings receive the same answer.

Fifth, the evaluation scores wobble and you cannot catch regressions. In evaluation, it is better to turn off drawing and fix it to greedy. What you want to measure is the change in the model, not the change in the random numbers.

What really matters in practice

What you will do in the next lab

You grow /root/work/tf-sample/sample.py one step at a time. You use only the standard library math and random — the system Python of this Pod has no numpy and it exists only inside /opt/onnx-lab/bin/python. You do not call a model either. You write down one set of logits yourself and build everything after that by hand.

You start with a stable softmax after dividing by the temperature and build up through entropy, top-k, top-p, a function that chains the four in a fixed order, inverse cumulative distribution sampling, and greedy and a temperature sweep. At the end you record in numbers whether drawing twice with the same seed gives the same sequence, and whether the lower-temperature side becomes the same as greedy.

The grader does not trust the explanations you wrote down. It actually imports your module, pokes at the functions with different logits and different temperatures every time, and checks them against probability vectors it computes separately. It does not judge random sampling itself — it looks only at the sequence of indices with the seed fixed and the probability vectors. That way, a correct implementation never fails by chance, and a wrong implementation never passes by chance.