Evaluation — Judge by Numbers, Not by Feel
In one line
You cannot judge the improvement of an LLM application without an evaluation set, and hallucination is not a bug you can eliminate but a risk you must manage.
Why this was needed
You fixed the prompt and it seems better. Is it really? If you judge after running it ten times, that is an experiment with a sample of ten, and you cannot tell whether the fixed prompt broke a different type of input.
Traditional software had tests. The expected output for an input was fixed, so you just compared whether they matched. An LLM's output differs a little every time, and there is more than one standard for what is "correct". So you have to design the evaluation method itself.
How it works
It is tidy to think of evaluation as three layers.
Layer 1: what can be checked deterministically. Is the output valid JSON, are the required fields present, does it keep within numeric ranges, does it contain forbidden words? These checks are cheap and reliable, so you automate them first. A good share of format errors are caught here.
Layer 2: what has a correct answer. Tasks whose expected output is fixed, such as classification or extraction, can be measured with accuracy and F1. The retrieval stage of RAG belongs here too. If you just have a list of correct documents, you get Recall@K and MRR. This is the layer where you can find room for improvement most cheaply.
Layer 3: what needs judgment. Is the answer faithful, is it helpful, is the tone appropriate? Human evaluation is the standard, but it is expensive and slow, so using a model as the judge is widely done. There are things to watch. Model judgment has a bias toward longer answers, and a tendency to be generous to answers it produced itself has also been reported. So check the correlation with human evaluation at least once, and it is more stable to break the judging criteria into a concrete checklist.
If you split hallucination into two kinds, the response differs.
- Extrinsic hallucination — it made up content that is not in the context. You can catch it with a faithfulness check, and reduce it by having the model mark supporting sentences or by making it refuse when below a threshold.
- Intrinsic hallucination — it summarized content that is in the context wrongly, or reversed it. This one is more dangerous, because the attached source actually earns it trust.
What it looks like in the field
An evaluation set does not need to be perfect. Even one with 50 items is far better than none. You can start by collecting the real user questions that failed. And if you add each case to the evaluation set whenever an incident happens, its value as a regression test grows over time.
Cost is also something to manage. Running the full evaluation on every commit is expensive, so you split it: always run the cheap layer 1 checks, and run the expensive layer 3 evaluation only before deployment.
Finally, the temperature setting is worth pointing out. When evaluating, lowering the temperature to reduce variation is better for comparison. But if it differs from the production setting you have to account for that difference, and if the temperature is high in production, it is more accurate to run several times and look at the variance too.
The order for reducing hallucination
The demand to "eliminate hallucination" is not solved by swapping the model. You reduce it through structure.
1. Give it grounding (RAG). Do not rely on what the model knows; put documents in with it. This alone greatly reduces factual errors. However, if retrieval is wrong, it becomes a grounded lie, so retrieval quality is answer quality.
2. Make it cite its grounding. If you make it attach a source piece number to every sentence, you can filter out sentences with no citation. Verification also becomes easier for people.
{"answer": "환불은 7영업일 안에 가능합니다 [2].",
"citations": [{"id": 2, "quote": "환불 요청은 결제일로부터 7영업일…"}]}
3. Make it say it does not know when it does not. State in the system prompt: "if it is not in the provided documents, answer 'this cannot be confirmed from the documents'". If you leave this out, the model always makes something up.
4. Attach a verifier. Check with a program whether the answer's citations actually exist in the documents. A citation that does not exist is itself evidence of hallucination.
How to measure hallucination
| Metric | How to measure | Automation |
|---|---|---|
| Citation accuracy | Is the cited sentence in the original? | Possible by string matching |
| Grounding faithfulness | Are the answer's claims supported by the documents? | LLM judge |
| No-answer rate | The fraction that said it does not know | Automatic |
| False no-answer | The fraction that said it does not know when it could answer | Automatic, since the evaluation set has the answers |
The last row matters. If you push it too hard to say it does not know, it evades even what it could answer. You have to look at both metrics together to keep the balance.
What to always put in an evaluation set
- Questions whose answer is in the documents — does it answer exactly?
- Questions whose answer is not in the documents — does it say it does not know, or make something up?
- Questions with a false premise — like "that policy that was abolished last year". Does it ask back about the premise?
- Questions that need several documents to be combined — does it avoid answering from a single piece?
- Questions that need freshness — does it avoid answering confidently on the basis of an old document?
The second and third show hallucination best. If you collect only the questions that go well, the score is high, but you end up with a system that keeps making things up in reality.
What to check in the following quiz
It checks whether you can judge what to measure at which layer, and how to handle the bias of model judgment.