AI Agents — A Graph, Not a Model
Same Score, Different Cases Failing
In one line
After you fix an agent, you must not judge whether it got better by a single score. Even if the score is the same, if one thing was fixed and another broken inside, it did not get better; it changed.
Why this was needed
We added "return" to the inquiry classification rules. The aim was to make "I want to return it", which had been falling into the "other" category, land in refund. When we measured with the golden set, the accuracy stayed the same. 83 percent to 83 percent.
We shipped it thinking "if it stayed the same, at least it did not get worse". Two days later, inquiries like "is the shipping fee refundable?" had all gone to the delivery team.
While fixing the rules we also changed the order of the checks, so one thing was fixed and one thing was broken. Five out of six are right, same as before, but the five that were right were different.
The score hides that fact. A score is a count, and what we needed to know was a name.
How it works
To do evaluation properly, you have to look at three things separately.
- Quality — how many did it get right on the golden set, and which ones did it get wrong.
- Cost — how many times did nodes run to handle one case.
- Path — did the path the same input takes change.
All three can be measured from the same run. You only need to wrap the nodes one layer.
def instrument(name, fn):
def wrapped(state):
started = time.perf_counter()
update = fn(state)
row = PROFILE.setdefault(name, {"calls": 0, "written": 0, "ms": 0.0})
row["calls"] += 1
row["written"] += len(update or {})
row["ms"] += (time.perf_counter() - started) * 1000.0
return update
return wrapped
There is one important choice here. We record time too, but we do not use it for judgment. Even with the same code and the same input, time varies with the machine and the load at that moment. If you judge "it got slower" by time, the regression check wobbles, and a wobbling check soon gets switched off. For judgment, use counts and sizes. Time is used as a clue when a person looks into it.
A golden set can be small. But it must not change
A golden set is "a bundle of inputs with answers attached by hand". Even six cases are useful — the condition is that it measures the same thing every time.
So fix the golden set in code or in a file, and when you change it, record the fact of changing it itself. If you draw a fresh sample on every run, you cannot compare yesterday and today.
One more thing. Leave cases that get it wrong in the golden set. A golden set where everything is right says only "it is going well now" and catches nothing. Only if it contains cases that actually fail can you see what was fixed when you fix something.
Look at regression as sets, not as scores
When comparing two versions, these two are what you must look at.
- Newly broken — something that was right in the old version is wrong in the new version
- Newly fixed — something that was wrong in the old version is right in the new version
If the two are equal in number, the score stays the same. But the weight of the two is usually not the same. Something that was already working getting broken hurts much more — because users are already relying on that behavior. So the order that is safe in practice is to first ask "is newly broken 0?" and then ask "is there anything newly fixed?"
Cost and path regress too
Even if quality stays the same, cost can rise. If you add one more branch, a node runs one more time, and if that node calls something external, that much money goes out. So keep the number of node calls as a regression metric too. Unlike time, this number is always the same for the same input.
The same goes for the path. Even if the same label comes out, if it came out by a different road, that is different behavior. The answer looks the same now, but at the next change it will split differently. Comparing the footprints (trace) shows this.
What it looks like in the field
First, recording only the score. This is the incident above. If you do not record which cases were wrong, you cannot retrace.
Second, the golden set changes with every run. If you measure with a random sample, the numbers differ every time, and you cannot tell whether it is a regression or the sample.
Third, judging performance regression by time. The check wobbles, and a wobbling check is ignored and in the end switched off.
Fourth, trying to attach observation later. The place to wrap nodes is cheapest when you build the graph. If you attach it later, you have to fix every node, and if you miss one, only that node does not appear in the ledger.
What really matters in practice
- Leave the names of the wrong cases together with the score. It is the only record you can retrace.
- Split regression into newly broken and newly fixed. Look at the broken side first.
- Use counts and sizes for judgment, and time only as a clue.
- Cost and path are regression metrics too. If the answer is the same but the road changed, it changed.
What you will do in the next lab
You grow /root/work/ageval/evalkit.py one step at a time. First you build a graph that classifies inquiries, wrap the nodes, and leave in the ledger the number of calls and the number of keys written. You pull out which node is busiest, and get the accuracy and the list of wrong cases from a golden set with answers attached by hand. Then you build a second version with the rules fixed and compare the two versions, and separate by name which cases broke and which were fixed even though the score is the same. Finally you measure the changes in call counts and paths too and leave them as a record.