TT Lab
Get started
Learn Learning paths Courses

LLM Engineering

Evaluation — Judge by Numbers, Not by Feel

Continue in TT Lab

In one line

You cannot judge the improvement of an LLM application without an evaluation set, and hallucination is not a bug you can eliminate but a risk you must manage.

Why this was needed

You fixed the prompt and it seems better. Is it really? If you judge after running it ten times, that is an experiment with a sample of ten, and you cannot tell whether the fixed prompt broke a different type of input.

Traditional software had tests. The expected output for an input was fixed, so you just compared whether they matched. An LLM's output differs a little every time, and there is more than one standard for what is "correct". So you have to design the evaluation method itself.

How it works

It is tidy to think of evaluation as three layers.

Layer 1: what can be checked deterministically. Is the output valid JSON, are the required fields present, does it keep within numeric ranges, does it contain forbidden words? These checks are cheap and reliable, so you automate them first. A good share of format errors are caught here.

Layer 2: what has a correct answer. Tasks whose expected output is fixed, such as classification or extraction, can be measured with accuracy and F1. The retrieval stage of RAG belongs here too. If you just have a list of correct documents, you get Recall@K and MRR. This is the layer where you can find room for improvement most cheaply.

Layer 3: what needs judgment. Is the answer faithful, is it helpful, is the tone appropriate? Human evaluation is the standard, but it is expensive and slow, so using a model as the judge is widely done. There are things to watch. Model judgment has a bias toward longer answers, and a tendency to be generous to answers it produced itself has also been reported. So check the correlation with human evaluation at least once, and it is more stable to break the judging criteria into a concrete checklist.

If you split hallucination into two kinds, the response differs.

What it looks like in the field

An evaluation set does not need to be perfect. Even one with 50 items is far better than none. You can start by collecting the real user questions that failed. And if you add each case to the evaluation set whenever an incident happens, its value as a regression test grows over time.

Cost is also something to manage. Running the full evaluation on every commit is expensive, so you split it: always run the cheap layer 1 checks, and run the expensive layer 3 evaluation only before deployment.

Finally, the temperature setting is worth pointing out. When evaluating, lowering the temperature to reduce variation is better for comparison. But if it differs from the production setting you have to account for that difference, and if the temperature is high in production, it is more accurate to run several times and look at the variance too.

The order for reducing hallucination

The demand to "eliminate hallucination" is not solved by swapping the model. You reduce it through structure.

1. Give it grounding (RAG). Do not rely on what the model knows; put documents in with it. This alone greatly reduces factual errors. However, if retrieval is wrong, it becomes a grounded lie, so retrieval quality is answer quality.

2. Make it cite its grounding. If you make it attach a source piece number to every sentence, you can filter out sentences with no citation. Verification also becomes easier for people.

{"answer": "환불은 7영업일 안에 가능합니다 [2].",
 "citations": [{"id": 2, "quote": "환불 요청은 결제일로부터 7영업일…"}]}

3. Make it say it does not know when it does not. State in the system prompt: "if it is not in the provided documents, answer 'this cannot be confirmed from the documents'". If you leave this out, the model always makes something up.

4. Attach a verifier. Check with a program whether the answer's citations actually exist in the documents. A citation that does not exist is itself evidence of hallucination.

How to measure hallucination

Metric How to measure Automation
Citation accuracy Is the cited sentence in the original? Possible by string matching
Grounding faithfulness Are the answer's claims supported by the documents? LLM judge
No-answer rate The fraction that said it does not know Automatic
False no-answer The fraction that said it does not know when it could answer Automatic, since the evaluation set has the answers

The last row matters. If you push it too hard to say it does not know, it evades even what it could answer. You have to look at both metrics together to keep the balance.

What to always put in an evaluation set

The second and third show hallucination best. If you collect only the questions that go well, the score is high, but you end up with a system that keeps making things up in reality.

What to check in the following quiz

It checks whether you can judge what to measure at which layer, and how to handle the bias of model judgment.