AI Agents — A Graph, Not a Model
Accuracy Held, but Other Cases Broke
Goal
You wrap the nodes to leave in a ledger what was done where, and get the accuracy and the list of wrong cases from a golden set. You compare against a second version with the rules fixed, separate newly broken and newly fixed cases by name, and measure even the changes in cost and path.
Why it matters
After fixing an agent, it is common to ask "did it get better?" with a single score. But even if the score is the same, if one thing was fixed and one broken inside, it did not get better; it changed. A score is a count, and what we need to know is a name. For observation, you only need to wrap the nodes one layer. There is one important choice here — we record time too, but we do not use it for judgment. Even with the same code and the same input, time varies with the machine and load, and a regression check that judges by time wobbles and eventually gets switched off. For judgment, use counts and sizes. A golden set may be small but must be the same every time. And you must leave the cases that fail now, so that you can see what was fixed when you fix something. A golden set where everything is right catches nothing. Even when quality stays the same, cost and path can regress. The number of node calls is always the same value for the same input, so it can be used as a regression metric, and if the footprints changed, it is different behavior even if the answer looks the same. The grader does not trust the explanations you wrote. It actually imports your module, runs it with arbitrary inquiries, and compares the golden set's verdicts and the differences between the two versions with values the grader computes separately.
Steps
- In /root/work/ageval/evalkit.py, create
LABELS,ANSWER,classify_v1,State, five nodes (normalize,refund,shipping,other,finish),ROUTE,build_graph(version="v1")andrun_one(text, version="v1"). - Add
PROFILE,reset_profile()andinstrument(name, fn), and makebuild_graphadd every node wrapped. - Add
hotspots(texts, version="v1")so that it gives the number of calls per node and the busiest node. - Add
GOLDENandevaluate(version="v1")so that they give the accuracy and the list of wrong cases. - Add
classify_v2andcompare(old="v1", new="v2")to separate the newly broken cases from the newly fixed cases. - Add
cost_delta(old="v1", new="v2")to measure the change in node call counts. - Add
path_changes(old="v1", new="v2")to find the cases whose footprints changed. - Record what you measured in /root/work/ageval/eval_report.json and /root/work/ageval/eval_report.md.
Notes
- Execution contract: the grader imports
/root/work/ageval/evalkit.pyas a Python module and uses the names listed above directly. It is not run as a script. classify_v1(text): if"환불"is present,"환불"; otherwise, if"배송"is present,"배송"; otherwise"기타"(the Korean words mean "refund", "delivery" and "other").classify_v2(text): if"배송"is present,"배송"; otherwise, if"환불"or"반품"is present,"환불"; otherwise"기타"(the second Korean word means "return"). The check order differs from v1 — that is the material for this lab.- State keys:
text,clean,label,answerandtrace. Onlytraceuses the appending reducer. A node name and a state key must not overlap — if they overlap, compiling givesValueError: 'x' is already being used as a state key. normalizecollapses the whitespace intextto a single space and puts it inclean. Classification looks atclean.refund,shippingandotherputlabelandANSWER[라벨](the placeholder stands for the label) and all three converge onfinish. All five nodes write their own names totrace.- Use
ROUTE = {"환불": "refund", "배송": "shipping", "기타": "other"}as the path map of the conditional edge. - Answer of
run_one(text, version):{"label": 문자열, "answer": 문자열, "trace": [...]}(the placeholders stand for strings). PROFILEis{노드이름: {"calls": 정수, "written": 정수, "ms": 실수}}(node name mapped to an integer call count, an integer count of written keys and a float of milliseconds).writtenis the sum of the number of keys in the dictionaries that node returned.msis recorded only and not used for judgment.- Answer of
hotspots(texts, version):{"calls": {...}, "total_calls": 정수, "busiest": 문자열, "timed": [...정렬된 노드 이름]}(an integer, a string, and the sorted node names). Each time you call it, you start a fresh ledger.busiestis the node with the most calls, and by name ascending on ties. GOLDENis six pairs of(문의, 정답라벨)(inquiry, correct label):("환불 절차 알려 주세요", "환불")(tell me the refund procedure),("반품하고 싶어요", "환불")(I want to return it),("배송비 환불되나요", "환불")(is the shipping fee refundable),("배송 언제 오나요", "배송")(when will the delivery arrive),("영수증 좀 보내 주세요", "기타")(please send me the receipt),("반품 배송비는 누가 내나요", "배송")(who pays the return shipping fee).- Answer of
evaluate(version):{"total": 정수, "correct": 정수, "accuracy": 실수, "wrong": [{"text":…, "gold":…, "got":…}, …]}(integers and a float).wrongfollows the order of the golden set. - Answer of
compare(old, new):{"old_accuracy": 실수, "new_accuracy": 실수, "newly_broken": [...정렬됨], "newly_fixed": [...정렬됨]}(floats and sorted lists). - Answer of
cost_delta(old, new):{"old_calls": 정수, "new_calls": 정수, "delta": 정수}(integers). It is the number of calls after running the whole golden set once each. - Answer of
path_changes(old, new):[{"text":…, "old": [...], "new": [...]}, …]intextascending order. Cases whose footprints are the same are not included. - This Pod has no internet. langgraph 0.2.60 is already installed.
- Official docs: Graph API overview · Use the graph API · Streaming
- Common mistakes: not clearing the ledger on each call (the numbers of the previous run get mixed in), judging regression by time, not leaving the names of wrong cases, and drawing the golden set fresh on every run.
Send inquiries down branches
In /root/work/ageval/evalkit.py, create LABELS, ANSWER, classify_v1, State, five nodes, ROUTE, build_graph(version="v1") and run_one(text, version="v1").
Classify by looking at the clean that normalize made. The condition function returns the label and ROUTE decides which node to go to — not making the label names and node names the same is meant to show clearly what the path map does. For now, the version argument accepts only "v1".
Wrap the nodes to leave a ledger
Add PROFILE, reset_profile() and instrument(name, fn), and make build_graph add all five nodes wrapped. In the ledger, leave calls, written and ms.
The wrapping function calls the original node, returns its result as it is, and writes numbers alongside. written is the number of keys in the dictionary that node returned. Record time too, but write in a comment that it is not used for judgment — why that is the heart of this step.
Which node is the busiest
Add hotspots(texts, version="v1") so that it gives {"calls": {...}, "total_calls": 정수, "busiest": 문자열, "timed": [...]} (an integer, a string and a list). Each time you call it, you start a fresh ledger.
If you do not clear the ledger, the numbers of the previous run get mixed in and you cannot tell what you are measuring. busiest is the node with the most calls; break ties by name ascending so that the same answer comes out every time. timed is the list of names of nodes whose time was recorded.
Six cases with answers attached by hand
Add the six pairs of GOLDEN and evaluate(version="v1") so that they give {"total", "correct", "accuracy", "wrong"}. wrong holds text, gold and got.
A golden set may be small, but it must be the same every time. And you must leave the cases that fail now, so that you can see what was fixed when you fix something — a golden set where everything is right catches nothing. Do not give only the accuracy; give the names of the wrong cases together with it.
The score is the same but different cases are wrong
Add classify_v2 and compare(old="v1", new="v2"). The answer is {"old_accuracy", "new_accuracy", "newly_broken", "newly_fixed"} and you sort the two lists.
classify_v2 is a version that makes "반품" (the Korean word for "return") count as refund while changing the check order. First measure the accuracy of the two versions, and ask whether those numbers alone let you judge. If you gather separately the cases that were right in the old version and are wrong in the new one, and the reverse, you can see what the score was hiding.
Even with the same quality, the cost differs
Add cost_delta(old="v1", new="v2") so that it gives {"old_calls", "new_calls", "delta"}. It is the number of node calls after running the whole golden set once each.
The number of calls is always the same value for the same input, so it can be used as a regression metric — that is how it differs from time. You can use the hotspots you built in the earlier step as it is. In this graph the number of nodes is the same whichever branch it goes to, so think about what delta tells you in that case.
The same answer came out by a different road
Add path_changes(old="v1", new="v2") so that it gives only the cases whose footprints changed, as [{"text", "old", "new"}, ...], in text ascending order.
Even if the label is the same, if the nodes passed through are different, it is different behavior. The answer looks the same now, but at the next change it will split differently. Do not include cases whose footprints are the same — a short list is one that people will look at.
Judge what to ship
Write old_accuracy, new_accuracy, newly_broken, newly_fixed, cost, path_changes, busiest and total_calls in /root/work/ageval/eval_report.json, and write /root/work/ageval/eval_report.md in four sections: ## 어디서 무엇을 했는지 어떻게 기록했나 (how you recorded what was done where), ## 점수만 보면 놓치는 것 (what you miss by looking only at the score), ## 비용과 경로도 회귀한다 (cost and path regress too) and ## 다음에 무엇을 하겠는가 (what you will do next).
cost is the answer of cost_delta as it is, and path_changes is the number of cases that changed. busiest and total_calls are the values when you run the whole golden set with v1. In the fourth section, write your own judgment of which version to ship after seeing these results — there is not just one right answer.