TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

Accuracy Held, but Other Cases Broke

Continue in TT Lab

Goal

You wrap the nodes to leave in a ledger what was done where, and get the accuracy and the list of wrong cases from a golden set. You compare against a second version with the rules fixed, separate newly broken and newly fixed cases by name, and measure even the changes in cost and path.

Why it matters

After fixing an agent, it is common to ask "did it get better?" with a single score. But even if the score is the same, if one thing was fixed and one broken inside, it did not get better; it changed. A score is a count, and what we need to know is a name. For observation, you only need to wrap the nodes one layer. There is one important choice here — we record time too, but we do not use it for judgment. Even with the same code and the same input, time varies with the machine and load, and a regression check that judges by time wobbles and eventually gets switched off. For judgment, use counts and sizes. A golden set may be small but must be the same every time. And you must leave the cases that fail now, so that you can see what was fixed when you fix something. A golden set where everything is right catches nothing. Even when quality stays the same, cost and path can regress. The number of node calls is always the same value for the same input, so it can be used as a regression metric, and if the footprints changed, it is different behavior even if the answer looks the same. The grader does not trust the explanations you wrote. It actually imports your module, runs it with arbitrary inquiries, and compares the golden set's verdicts and the differences between the two versions with values the grader computes separately.

Steps

  1. In /root/work/ageval/evalkit.py, create LABELS, ANSWER, classify_v1, State, five nodes (normalize, refund, shipping, other, finish), ROUTE, build_graph(version="v1") and run_one(text, version="v1").
  2. Add PROFILE, reset_profile() and instrument(name, fn), and make build_graph add every node wrapped.
  3. Add hotspots(texts, version="v1") so that it gives the number of calls per node and the busiest node.
  4. Add GOLDEN and evaluate(version="v1") so that they give the accuracy and the list of wrong cases.
  5. Add classify_v2 and compare(old="v1", new="v2") to separate the newly broken cases from the newly fixed cases.
  6. Add cost_delta(old="v1", new="v2") to measure the change in node call counts.
  7. Add path_changes(old="v1", new="v2") to find the cases whose footprints changed.
  8. Record what you measured in /root/work/ageval/eval_report.json and /root/work/ageval/eval_report.md.

Notes

Send inquiries down branches

In /root/work/ageval/evalkit.py, create LABELS, ANSWER, classify_v1, State, five nodes, ROUTE, build_graph(version="v1") and run_one(text, version="v1").

Classify by looking at the clean that normalize made. The condition function returns the label and ROUTE decides which node to go to — not making the label names and node names the same is meant to show clearly what the path map does. For now, the version argument accepts only "v1".

Wrap the nodes to leave a ledger

Add PROFILE, reset_profile() and instrument(name, fn), and make build_graph add all five nodes wrapped. In the ledger, leave calls, written and ms.

The wrapping function calls the original node, returns its result as it is, and writes numbers alongside. written is the number of keys in the dictionary that node returned. Record time too, but write in a comment that it is not used for judgment — why that is the heart of this step.

Which node is the busiest

Add hotspots(texts, version="v1") so that it gives {"calls": {...}, "total_calls": 정수, "busiest": 문자열, "timed": [...]} (an integer, a string and a list). Each time you call it, you start a fresh ledger.

If you do not clear the ledger, the numbers of the previous run get mixed in and you cannot tell what you are measuring. busiest is the node with the most calls; break ties by name ascending so that the same answer comes out every time. timed is the list of names of nodes whose time was recorded.

Six cases with answers attached by hand

Add the six pairs of GOLDEN and evaluate(version="v1") so that they give {"total", "correct", "accuracy", "wrong"}. wrong holds text, gold and got.

A golden set may be small, but it must be the same every time. And you must leave the cases that fail now, so that you can see what was fixed when you fix something — a golden set where everything is right catches nothing. Do not give only the accuracy; give the names of the wrong cases together with it.

The score is the same but different cases are wrong

Add classify_v2 and compare(old="v1", new="v2"). The answer is {"old_accuracy", "new_accuracy", "newly_broken", "newly_fixed"} and you sort the two lists.

classify_v2 is a version that makes "반품" (the Korean word for "return") count as refund while changing the check order. First measure the accuracy of the two versions, and ask whether those numbers alone let you judge. If you gather separately the cases that were right in the old version and are wrong in the new one, and the reverse, you can see what the score was hiding.

Even with the same quality, the cost differs

Add cost_delta(old="v1", new="v2") so that it gives {"old_calls", "new_calls", "delta"}. It is the number of node calls after running the whole golden set once each.

The number of calls is always the same value for the same input, so it can be used as a regression metric — that is how it differs from time. You can use the hotspots you built in the earlier step as it is. In this graph the number of nodes is the same whichever branch it goes to, so think about what delta tells you in that case.

The same answer came out by a different road

Add path_changes(old="v1", new="v2") so that it gives only the cases whose footprints changed, as [{"text", "old", "new"}, ...], in text ascending order.

Even if the label is the same, if the nodes passed through are different, it is different behavior. The answer looks the same now, but at the next change it will split differently. Do not include cases whose footprints are the same — a short list is one that people will look at.

Judge what to ship

Write old_accuracy, new_accuracy, newly_broken, newly_fixed, cost, path_changes, busiest and total_calls in /root/work/ageval/eval_report.json, and write /root/work/ageval/eval_report.md in four sections: ## 어디서 무엇을 했는지 어떻게 기록했나 (how you recorded what was done where), ## 점수만 보면 놓치는 것 (what you miss by looking only at the score), ## 비용과 경로도 회귀한다 (cost and path regress too) and ## 다음에 무엇을 하겠는가 (what you will do next).

cost is the answer of cost_delta as it is, and path_changes is the number of cases that changed. busiest and total_calls are the values when you run the whole golden set with v1. In the fourth section, write your own judgment of which version to ship after seeing these results — there is not just one right answer.