Building a Response-Quality Regression Harness
Goal
Build an evaluation harness from scratch that measures quality and latency together, and complete it up to a CI gate that automatically judges regressions against a baseline.
Why it matters
Changing a model or a prompt without an evaluation harness is the same as refactoring without tests. You decide with "it seems better" after asking a few questions, two weeks later you get an inquiry that quality dropped on a certain type, and you have no basis for deciding whether to roll back — and this repeats. LLM output is not deterministic, so perfect evaluation is impossible, but you can measure enough to catch regressions. What is especially important in this lab is step 5 — measuring only quality is half. If quality improved by 2 percentage points but latency doubled, that is not an improvement. And the CI gate in step 7 makes all of this actually work. An evaluation that a person has to remember to run eventually does not get run.
Steps
- Read
/opt/fixtures/llms/evalset.jsonl(40 cases, each line withid,prompt,expectedandgrader) and writecases=40 graders=<쉼표로 이은 종류>in/root/ev/loaded.txt(the placeholder is the grader types joined by commas). - Run all the cases against the server on 8170 with
/root/ev/run.pyand produce/root/ev/results.jsonl. Each line must haveid,outputandlatency_ms, and there must be 40 lines. - Grade the cases whose
graderisexactby exact match. Writecases=<n> passed=<n> score=<0~1 소수>in/root/ev/exact.txt(the placeholders are counts and a decimal between 0 and 1). - Grade the cases whose
graderiscontainsby keyword inclusion ratio. Write it in/root/ev/contains.txtin the same format, andscoremust be at least 0 and at most 1. - Set the latency SLO to 800ms and count the violations. Write
slo_ms=800 violations=<n> p99_ms=<수>in/root/ev/latency.txt(the placeholder is a number). - Save the overall score in
/root/ev/baseline.jsonas{"score":<수>,"p99_ms":<수>,"cases":40}(the placeholders are numbers). Then compare the new run result against the baseline with/root/ev/compare.pyand writebaseline=<수> current=<수> delta=<부호 있는 수> verdict=<PASS|FAIL>in/root/ev/regression.txt(the placeholders are numbers, and delta is a signed number). The threshold is -0.03. /root/ev/gate.shends with exit code 1 on a regression and 0 otherwise. Test both cases and writepass_exit=0 fail_exit=1in/root/ev/gate.out.
Notes
- The backend at 127.0.0.1:8170 is not pre-started in this Pod. It only needs to take
{"prompt":..., "max_tokens":n}atPOST /generateand return{"text":..., "usage":{...}}, so start by using the server you built in the earlier token streaming lab as it is, or by starting a minimal server with the same contract yourself. To create a regression, it is convenient to have one knob with which you can change the quality from outside. - The evaluation set must reflect the real traffic distribution. If you gather only the cases that work well, you cannot catch regressions.
- The threshold must be larger than the noise. 3 percentage points is the starting point, and it must be larger if the evaluation set is small.
- If you build a loop that adds failure cases discovered through user inquiries to the evaluation set, the same failure will not happen twice.
- Common mistake 1: measuring only quality and leaving out latency and cost.
- Common mistake 2: building the harness and not putting it in CI — an evaluation that a person has to remember to run eventually does not get run.
Load the evaluation set
Read /opt/fixtures/llms/evalset.jsonl (40 cases, each line with id, prompt, expected and grader) and write cases=40 graders=<쉼표로 이은 종류> in /root/ev/loaded.txt (the placeholder is the grader types joined by commas).
One JSONL line is one case. The grading method may differ for each case.
Run all the cases with a runner
Run all the cases against the server on 8170 with /root/ev/run.py and produce /root/ev/results.jsonl. Each line must have id, output and latency_ms, and there must be 40 lines.
You must keep the results together with the originals so that you can later see what went wrong.
Grade by exact match
Grade the cases whose grader is exact by exact match. Write cases=<n> passed=<n> score=<0~1 소수> in /root/ev/exact.txt (the placeholders are counts and a decimal between 0 and 1).
It is the method that suits structured output. Normalize whitespace and case first.
Grade with partial credit
Grade the cases whose grader is contains by keyword inclusion ratio. Write it in /root/ev/contains.txt in the same format, and score must be at least 0 and at most 1.
For free-form prose, you score by the keyword inclusion ratio. A value between 0 and 1 must come out.
Count latency SLO violations
Set the latency SLO to 800ms and count the violations. Write slo_ms=800 violations=<n> p99_ms=<수> in /root/ev/latency.txt (the placeholder is a number).
Measuring only quality is half. Getting slower is also a regression.
Save the baseline and judge regressions
Save the overall score in /root/ev/baseline.json as {"score":<수>,"p99_ms":<수>,"cases":40} (the placeholders are numbers). Then compare the new run result against the baseline with /root/ev/compare.py and write baseline=<수> current=<수> delta=<부호 있는 수> verdict=<PASS|FAIL> in /root/ev/regression.txt (the placeholders are numbers, and delta is a signed number). The threshold is -0.03.
Without something to compare against, you cannot judge. The threshold must be larger than the noise.
Build a CI gate script
/root/ev/gate.sh ends with exit code 1 on a regression and 0 otherwise. Test both cases and write pass_exit=0 fail_exit=1 in /root/ev/gate.out.
The goal is to make it so that no person has to remember. Report whether it passed through the exit code.