TT Lab
Get started
Learn Learning paths Courses

LLM Serving

What Happens When You Change Models on a Hunch

Continue in TT Lab

In one line

Changing a model or a prompt without an evaluation harness is the same as refactoring without tests. Feeling that it got better and actually getting better are different things.

Why this was needed

A new model comes out. You ask a few questions and the answers are better. You switch. Two weeks later an inquiry comes in that quality dropped on a certain type of question. You have to decide whether to roll back, but you have no basis.

The reason this situation repeats is that LLM output is not deterministic and the evaluation criteria are subjective. In many cases you cannot expect "exactly this value" as in a unit test. On top of that, human memory is swayed greatly by the last few cases. One impressive failure overturns the judgment of overall quality.

Still, you can measure. Even if not perfect, it is enough to catch regressions.

The four parts of an evaluation harness

The evaluation set. Representative questions drawn from real traffic and their expected results. 30–100 cases is enough to start detecting regressions. What matters is reflecting the real distribution — if you gather only the cases that work well, you cannot catch regressions. Match the ratio of traffic types and include at least 20% known-difficult cases.

The grader. It depends on the type.

Output type Grading method Caution
JSON, classification, extraction Exact match Normalize field order and whitespace before comparing
Short factual answer Keyword inclusion, regex Allow synonyms as a list
Long prose LLM judge Pin the versions of the judge model and prompt
Code Run and pass tests The most trustworthy

An LLM judge has biases. It scores long answers higher than short ones, prefers the style of models from its own family, and prefers the candidate shown first. So when having it do an A/B comparison, ask twice with the order randomly swapped, and if the result flips, count it as a tie.

The baseline. Save the score of the combination currently in operation. Without this there is nothing to compare against. You must record the model, prompt, temperature and evaluation set version together to reproduce it later.

The gate. If the score of the new combination is lower than the baseline by more than a certain margin, it is judged a failure.

How to set the threshold — measure the noise first

You must not set the threshold by gut feeling. Run the same settings three times, measure the width of the wobble first, and use a value larger than that as the threshold.

# 같은 모델·프롬프트로 3회
run 1: 0.847   run 2: 0.861   run 3: 0.839
→ 노이즈 폭 약 2.2%p → 임계는 3~4%p 로 잡는다

If the evaluation set is small, the noise is large. In a 50-case evaluation set, one case flipping moves it by 2 percentage points. Statistically speaking, the 95% confidence interval of a 0.85 observed in 50 cases is roughly 0.75–0.92. You must not judge by a 1–2 percentage point difference on a small evaluation set.

Even setting the temperature to 0 is not fully deterministic. The results can vary depending on the batch size and the order of floating-point accumulation, so the same input sometimes produces different outputs.

Measuring only quality is half

You must measure latency and cost in the same harness along with quality.

Axis Value measured Gate example
Quality Evaluation set score Fail if baseline −3 percentage points or lower
Latency p50, p95, time to first token Fail if p95 exceeds 1.5 times the baseline
Cost Token cost per 1,000 cases A person approves if it exceeds 2 times the baseline

You must look at p95, not the average. There really are changes where the average latency is good but p95 triples. Users do not experience the average; they experience the latency of their own request.

Time to first token (TTFT) is also measured separately. In a streaming UI, TTFT, more than the total completion time, decides the perceived quality.

What it looks like in the field

Putting the harness into CI is decisive. If it runs automatically on prompt-change PRs and a regression blocks the merge, no person has to remember. This is the actual meaning of treating prompts like code.

The evaluation set must be alive. If you build a loop that adds failure cases discovered through user inquiries to the evaluation set, the same failure will not happen twice. It is the same principle as regression tests.

But there is one pitfall. If you keep fixing the prompt to fit the evaluation set, you overfit to the evaluation set. The score goes up while real traffic gets worse. To prevent this, split the evaluation set in two, separating a development set (the one you look at repeatedly) from a validation set (the one you look at only occasionally). If the validation score stops following the development score, that is a sign of overfitting.

What to put in the evaluation set

The most common mistake in building an evaluation set is gathering only what works well. Then the score is always high and regressions are not caught. You deliberately mix five kinds.

  1. Representative cases — the type that is most common in real traffic. About half of the total.
  2. Edge cases — when the input is empty, very long, or has another language mixed in.
  3. Known failures — ones that came in as inquiries in the past. This is the core of regression detection.
  4. Things that must be refused — questions that must not be answered. You check that the model has become well-behaved.
  5. Traps — questions with a wrong premise ("that person who got the 2026 Nobel Prize, you know"). You see whether the model asks back about the premise or makes something up.

When writing the expected result, write it not as a single correct answer but as an allowed set.

{"id": "refund-01",
 "input": "환불 언제까지 돼요?",
 "must_include": ["7일", "영업일"],
 "must_not_include": ["환불 불가", "죄송"],
 "max_tokens": 200}

must_not_include is surprisingly useful. A considerable share of quality incidents are "something that should not be there came out", not "something that should be there is missing".

What you will do in the next lab

You load the evaluation set and build a runner, implement an exact-match grader and a partial-credit grader, count latency SLO violations, run the same settings three times to measure the noise width, save the baseline and then judge regressions, and finally turn it into a CI gate script.