TT Lab
Get started
Learn Learning paths Courses

LLM Serving

Building a Response-Quality Regression Harness

Continue in TT Lab

Goal

Build an evaluation harness from scratch that measures quality and latency together, and complete it up to a CI gate that automatically judges regressions against a baseline.

Why it matters

Changing a model or a prompt without an evaluation harness is the same as refactoring without tests. You decide with "it seems better" after asking a few questions, two weeks later you get an inquiry that quality dropped on a certain type, and you have no basis for deciding whether to roll back — and this repeats. LLM output is not deterministic, so perfect evaluation is impossible, but you can measure enough to catch regressions. What is especially important in this lab is step 5 — measuring only quality is half. If quality improved by 2 percentage points but latency doubled, that is not an improvement. And the CI gate in step 7 makes all of this actually work. An evaluation that a person has to remember to run eventually does not get run.

Steps

  1. Read /opt/fixtures/llms/evalset.jsonl (40 cases, each line with id, prompt, expected and grader) and write cases=40 graders=<쉼표로 이은 종류> in /root/ev/loaded.txt (the placeholder is the grader types joined by commas).
  2. Run all the cases against the server on 8170 with /root/ev/run.py and produce /root/ev/results.jsonl. Each line must have id, output and latency_ms, and there must be 40 lines.
  3. Grade the cases whose grader is exact by exact match. Write cases=<n> passed=<n> score=<0~1 소수> in /root/ev/exact.txt (the placeholders are counts and a decimal between 0 and 1).
  4. Grade the cases whose grader is contains by keyword inclusion ratio. Write it in /root/ev/contains.txt in the same format, and score must be at least 0 and at most 1.
  5. Set the latency SLO to 800ms and count the violations. Write slo_ms=800 violations=<n> p99_ms=<수> in /root/ev/latency.txt (the placeholder is a number).
  6. Save the overall score in /root/ev/baseline.json as {"score":<수>,"p99_ms":<수>,"cases":40} (the placeholders are numbers). Then compare the new run result against the baseline with /root/ev/compare.py and write baseline=<수> current=<수> delta=<부호 있는 수> verdict=<PASS|FAIL> in /root/ev/regression.txt (the placeholders are numbers, and delta is a signed number). The threshold is -0.03.
  7. /root/ev/gate.sh ends with exit code 1 on a regression and 0 otherwise. Test both cases and write pass_exit=0 fail_exit=1 in /root/ev/gate.out.

Notes

Load the evaluation set

Read /opt/fixtures/llms/evalset.jsonl (40 cases, each line with id, prompt, expected and grader) and write cases=40 graders=<쉼표로 이은 종류> in /root/ev/loaded.txt (the placeholder is the grader types joined by commas).

One JSONL line is one case. The grading method may differ for each case.

Run all the cases with a runner

Run all the cases against the server on 8170 with /root/ev/run.py and produce /root/ev/results.jsonl. Each line must have id, output and latency_ms, and there must be 40 lines.

You must keep the results together with the originals so that you can later see what went wrong.

Grade by exact match

Grade the cases whose grader is exact by exact match. Write cases=<n> passed=<n> score=<0~1 소수> in /root/ev/exact.txt (the placeholders are counts and a decimal between 0 and 1).

It is the method that suits structured output. Normalize whitespace and case first.

Grade with partial credit

Grade the cases whose grader is contains by keyword inclusion ratio. Write it in /root/ev/contains.txt in the same format, and score must be at least 0 and at most 1.

For free-form prose, you score by the keyword inclusion ratio. A value between 0 and 1 must come out.

Count latency SLO violations

Set the latency SLO to 800ms and count the violations. Write slo_ms=800 violations=<n> p99_ms=<수> in /root/ev/latency.txt (the placeholder is a number).

Measuring only quality is half. Getting slower is also a regression.

Save the baseline and judge regressions

Save the overall score in /root/ev/baseline.json as {"score":<수>,"p99_ms":<수>,"cases":40} (the placeholders are numbers). Then compare the new run result against the baseline with /root/ev/compare.py and write baseline=<수> current=<수> delta=<부호 있는 수> verdict=<PASS|FAIL> in /root/ev/regression.txt (the placeholders are numbers, and delta is a signed number). The threshold is -0.03.

Without something to compare against, you cannot judge. The threshold must be larger than the noise.

Build a CI gate script

/root/ev/gate.sh ends with exit code 1 on a regression and 0 otherwise. Test both cases and write pass_exit=0 fail_exit=1 in /root/ev/gate.out.

The goal is to make it so that no person has to remember. Report whether it passed through the exit code.