TT Lab
Get started
Learn Learning paths Courses

LLM Engineering

Prompts and Context — The Window Got Wider but Not Free

Continue in TT Lab

In one line

The context window is the upper limit on the number of tokens a model can see at once, and being able to put something in is a completely different matter from it being better to put it in.

Why this was needed

As context lengths grew, the thought "why not just put everything in?" became natural. In practice three costs come with it.

Money. Input tokens are billed too. If you put the whole document in every request, that cost is multiplied by the number of requests.

Latency. The longer the input, the longer the time to the first token. This is because the attention computation grows faster than linearly with length.

Quality. This is the most counterintuitive. The more irrelevant content you put in, the worse the ability to find the correct answer gets. It has been reported several times that information placed in the middle of a long context is especially easy to ignore. So the assumption that it is enough for the correct answer to be in there is dangerous. Where and how it is placed matters.

How it works

Here are a few practical principles of prompt design.

Role and instructions first, data after. The model reads the instructions first and interprets what follows within that frame. If you pour in the data first and attach the instructions at the end, the earlier part is read without direction.

Show the output format with an example. One actual JSON example is far stronger than the words "answer in JSON". However, examples eat tokens, so you need a balance between their number and length.

Have it break the work into steps. For complex reasoning, having the model write out the intermediate process raises accuracy. But if you force this on every question, tokens and latency grow, and on simple questions it can even produce confusing answers.

Give the context a priority. Instead of concatenating the retrieved documents as they are, a strategy of placing the highest-scoring ones at the front and back ends is commonly used. It turns the observation that the middle is weak to your advantage.

The system prompt is a contract. When the constraints written there are not obeyed, it is usually one of two things. Either the constraint is vague (for example, "be concise"), or the data that follows conflicts with the constraint. In particular, if the text a user inserted contains sentences that look like instructions, the model may try to follow them. The boundary that content coming from outside is data, not instructions has to be kept on both sides: in the prompt structure and in the application.

What it looks like in the field

It is good to manage the context budget explicitly. Decide in numbers how big the whole window is, how many tokens the system prompt takes, how many tokens you allot to retrieval results, and how many tokens you leave for the output. If you do not, you get the incident where the output is cut off on the day a user pastes in a long document.

And prompts need version control like code. When you change one line, you cannot tell without evaluation which cases got better and which got worse. Editing a prompt without an evaluation set is just gambling.

In a long context, the middle disappears

A context window of 128K does not mean everything inside it is used equally. Several studies have reported the same phenomenon: the very beginning and the very end are used well, and the middle is missed (lost in the middle).

[시스템 프롬프트] [문서 1] [문서 2] … [문서 20] [질문]
      ↑ 잘 본다              ↑ 여기가 약하다        ↑ 잘 본다

So when you put in retrieved documents, place the most relevant ones at the very beginning or the very end. Sorting by relevance and putting the best first, then repeating the question once more at the very end, works well in practice.

Often picking just 5 documents by reranking and putting those in gives higher accuracy than putting in 20. Putting in everything just because you can makes things worse.

A structure that saves tokens

There are a few ways to cut cost while getting the same result.

Treating prompts like code

prompts/
  summarize/
    v3.md            ← 버전을 파일로
    eval.jsonl       ← 그 버전의 평가셋
    baseline.json    ← 그 버전의 점수

If you bury a prompt as a string in code, there is no record of who changed it, when and why. If you keep it as a file and give it a version, rollback and A/B comparison become possible.

Also, have an evaluation harness run automatically on the pull request that changes a prompt. Without this, you end up changing things because it "feels better", and users discover the regression first.

What to check in the following quiz

It checks whether you can explain why making the context longer is not always a gain, and how the structure of a prompt changes the result.