AI Agents — A Graph, Not a Model
The One Call With Made-Up Arguments
Goal
You write in code the discipline of the side that calls tools. You keep a contract table in one place, measure the chosen arguments before calling, measure the returned result without trusting it, treat failure as a value rather than an exception and leave the reason in the state, fix the next attempt with that reason, and remember results so the same arguments are not called twice.
Why it matters
The place where an agent with tools attached first breaks is not the model but the call boundary. If the model makes up one argument name, a TypeError pops out of the graph, the whole run ends, and neither the state nor the path taken remains. Worse is the case where no exception is raised — the tool returned an empty list, and a sentence like "stock in 3 places" gets built on top of it.
So this lab starts by writing the contract as a value. If you tie the name, the argument schema, the shape of the return and the actual function into one table, you do not need measuring code per tool and the door for calling narrows to one place. In front of that door you measure the arguments, behind that door you measure the result, and in front of that door you put a memory.
In this lab, rules stand in for the place that picks the arguments. A real agent asks the model for the tool name and arguments in that place (the pick node). The rest of the structure is the same — if anything, the more the model picks, the more you need pre-call validation.
The grader does not trust the explanations you wrote. It actually imports your module, pokes at the validation functions with different values each time, and counts how many times the tool body ran to check "was it blocked before the call". The sku, region and count change with each run.
Steps
- In /root/work/agtool/tools.py, create the data (
REGIONS,STOCK),MAX_ROWSandCALLS, two tools (stock_lookup,legacy_stock) and the contract tableTOOLS. - Create
validate_args(name, args)to measure the arguments before calling. If it passes, it returns an empty string, and otherwise one reason. - Create
call_tool(name, args)so that only calls that pass validation are actually called. It does not let exceptions out and returns them as{"ok", "value", "error"}. - Create
validate_result(name, value)and makecall_toolmeasure the result too. The reasons areshape,empty,too_bigandrange. - Create
CACHEandcache_key(name, args)so that the same tool is not called twice with the same arguments. It remembers only successes. - Create
State, four nodes (pick,fetch,reply,giveup),build_naive()andrun_naive(order)to treat failure as a path and not an exception. - Add
MAX_ATTEMPTS,REPAIRS,can_repair, arepairnode,build_graph()andhandle(order)so that it reads the reason, fixes it and tries again. - Record what you confirmed in /root/work/agtool/tool_report.json and /root/work/agtool/tool_report.md.
Notes
- Execution contract: the grader imports
/root/work/agtool/tools.pyas a Python module and usesREGIONS,STOCK,MAX_ROWS,CALLS,TOOLS,stock_lookup,legacy_stock,validate_args,validate_result,call_tool,CACHE,cache_key,run_naive,MAX_ATTEMPTS,REPAIRS,can_repairandhandledirectly. It is not run as a script, soif __name__ == "__main__"is not needed. - The data for this lab is written like this.
REGIONS = {"seoul": ["gasan", "guro", "mapo"], "busan": ["sasang", "haeundae"], "jeju": ["hallim"]},STOCK = {"A-1001": {"gasan": 4, "guro": 2, "mapo": 7, "sasang": 1}, "A-1002": {"guro": 5, "haeundae": 3}, "B-2001": {"gasan": 9, "guro": 1, "mapo": 9, "hallim": 2}, "B-2002": {"sasang": 6, "haeundae": 6}},MAX_ROWS = 2. - The signatures of the two tools are the same,
(sku, region, limit), and what they return is also the same,{"sku": 문자열, "rows": [{"warehouse": 문자열, "count": 정수}, ...]}(the placeholders stand for a string and an integer).rowsis in count descending order, and by warehouse name ascending when equal.stock_lookupreturns only up tolimitrows, andlegacy_stockignoreslimitand returns all the warehouses in that region (imitating an old version). CALLSis a dictionary that counts, by name, how many times the tool body ran. Raise it by 1 on the first line of both tools. The grader uses this value to check "was it blocked before the call".- The shape of one
TOOLSentry:{"fn": 함수, "args": {인자이름: {"type": "str"|"int", "required": True, "allowed": [...], "min": N, "max": N}}, "returns": {"sku": {"type": "str"}, "rows": {"type": "list", "max_len": MAX_ROWS, "item": {"warehouse": {"type": "str"}, "count": {"type": "int", "min": 0}}}}}(the placeholders stand for the function and the argument name). Theallowedofregionis the keys ofREGIONS, andlimithasmin1 andmaxMAX_ROWS. - Reasons of
validate_args:unknown_tool·missing:<이름>·extra:<이름>·type:<이름>·allowed:<이름>·range:<이름>(the placeholder stands for the argument name). If it passes, it is an empty string. The grader throws only arguments with exactly one defect. - Reasons of
validate_result:shape(the shape or type differs) ·empty(rowsis empty) ·too_big(rowsis longer thanmax_len) ·range(countis outside the contract's range). If it passes, it is an empty string. - The
errorthatcall_toolreturns is prefixed with the source:args:<사유>for the argument side,result:<사유>for the result side, andraised:<예외이름>if the tool threw an exception (the placeholders stand for the reason and the exception name). If it succeeded, it is an empty string. - The
orderofhandle(order)is{"sku": ..., "region": ..., "limit": ..., "prefer": ...}. Ifpreferis there, it uses that tool, and otherwisestock_lookup. The answer is{"ok", "answer", "rows", "errors", "fixes", "attempts", "calls", "tool_runs"}, andtool_runsis the number of times the tool body ran while handling this case. REPAIRSis{"args:allowed:region": "region", "args:range:limit": "limit", "result:too_big": "tool"}. You fix the region toFALLBACK_REGION("seoul"), the count toMAX_ROWS, and the tool tostock_lookup.MAX_ATTEMPTSis 3.- Node names and state keys are in the same namespace. If you put a
toolkey in the state and also name a nodetool, compiling dies withValueError: 'tool' is already being used as a state key. - This Pod has no internet.
pip installdoes not work. langgraph 0.2.60 is already installed (python3 -c "import langgraph"). - Official docs: Workflows and agents · Use the graph API · Graph API overview
- Common mistakes: measuring the arguments after calling (the body has already run), accepting
Trueas a count becauseisinstance(True, int)is true, passing an empty result along as a success, remembering failures too so you cannot call again after fixing, and fixing and fixing again with no attempt limit.
Write the contract in one place as a value
In /root/work/agtool/tools.py, create REGIONS, STOCK, MAX_ROWS and CALLS, two tools (stock_lookup, legacy_stock) and the contract table TOOLS. One TOOLS entry holds three things, fn, args and returns, and both tools raise CALLS by 1 on the first line.
args is a dictionary that writes, for each argument name, type and required and, if needed, allowed, min and max. In returns, write the type of the keys returned and the max_len and item of rows. stock_lookup returns only up to limit rows, while legacy_stock returns everything, ignoring limit — you will see in a later step why result validation is needed, using this tool. Use the data table in the notes section as it is.
Measure the arguments before calling
Create validate_args(name, args). Reading the contract table, measure missing arguments, arguments not in the contract, type, allowlist and range, and return an empty string if it passes, and otherwise one reason (like missing:limit).
Use the reason names exactly as they are in the notes section. In Python, bool is a subtype of int, so isinstance(True, int) is true — if True arrives in the count slot, you must block it with type:. If the tool name is unknown, it is unknown_tool. What to return first when there are several defects is up to you, but the grader throws only arguments with exactly one defect.
Do not call with wrong arguments
Create call_tool(name, args). Only calls that pass validation are actually called, and exceptions are not let out. The answer is always {"ok": 참거짓, "value": 결과 또는 None, "error": 사유} (a boolean, the result or None, and the reason), and you prefix args: to argument-side reasons and raised: to exceptions the tool threw.
Take the actual function out of TOOLS[name]["fn"] and call it with fn(**args) — narrowing the door for calling to one place is the point. If it is caught in validation, end right there. The grader checks whether CALLS stays the same after giving wrong arguments. The approach of catching the exception after calling does not pass this check.
Do not trust what was returned
Create validate_result(name, value) and make call_tool measure the result too. The reasons are shape, empty, too_big and range, and when one is caught, call_tool returns it with result: prefixed to error.
Read the contract's returns and measure — if you write it by hand per tool, you miss it when tools are added. If rows is empty, it is empty (if you pass it along as a success because it is not an error, you end up making up an answer). If it is longer than max_len, it is too_big, and legacy_stock produces exactly that case. If a count inside a row is smaller than the contract's min, it is range.
Do not ask the same thing twice
Create CACHE and cache_key(name, args) so that call_tool remembers the result of the same tool with the same arguments. The same key must come out even if the arguments were written in a different order, and it remembers only successes.
If you join the name with json.dumps(args, sort_keys=True, ensure_ascii=False), you get a key that does not wobble with order. Look up the memory after passing validation, right before calling. If you remember failures too, the path to fix and call again is blocked, so put in only successes. The grader calls twice with the same arguments to see whether CALLS goes up only once, and whether the body runs both times for a failed call.
Treat failure as a path
Create State, four nodes (pick, fetch, reply, giveup), build_naive() and run_naive(order). fetch does not raise an exception even when it fails and leaves the reason in errors, and if it fails it goes to giveup. run_naive returns {"ok", "answer", "rows", "errors", "calls"}.
pick chooses the tool name and arguments from the order — a real agent asks the model in this place. Attach an appending reducer to errors and calls, and an adding reducer to attempts. The version in this step does not fix — it calls once and, if it fails, gives up as is. reply builds the answer sentence from the first row of rows, and giveup writes the last reason into the answer.
Read the reason, fix it and try again
Add MAX_ATTEMPTS = 3, FALLBACK_REGION, REPAIRS, can_repair(reason), a repair node, build_graph() and handle(order). If the reason is fixable, fix the arguments or the tool and go back to fetch, and if it cannot be fixed or it reaches the limit, go to giveup.
repair reads the last reason in errors, fixes it according to the REPAIRS table, and leaves what it fixed in fixes. It can be used in this place thanks to having left the reason in the state — if you had thrown an exception, there would be nothing but 'something went wrong'. An empty result is still empty if you ask again, so it is on the side that cannot be fixed. Get tool_runs of handle by subtracting the sum of CALLS before and after the call.
Record what you measured
Write tools, blocked_before_call, bad_result_reasons, memo_second_call_runs, repaired, attempts and gave_up in /root/work/agtool/tool_report.json, and write /root/work/agtool/tool_report.md in four sections: ## 도구 계약을 어디에 적었나 (where you wrote the tool contract), ## 부르기 전에 무엇을 막았나 (what you blocked before calling), ## 돌려준 결과를 어떻게 믿지 않았나 (how you did not trust the returned result) and ## 실패를 어떤 경로로 다뤘나 (by what path you handled failure).
Do not write the numbers by hand; get them by actually running your module. blocked_before_call is the number of times the tool body ran after you tried calling with wrong arguments, bad_result_reasons is the sorted reasons validate_result gave for the four bad results, and memo_second_call_runs is the number of times the body ran when you called a second time with the same arguments. repaired and attempts are obtained by passing an order with a wrong region to handle, and gave_up is the inverted ok of an order that yields an empty result. In the section ## 실패를 어떤 경로로 다뤘나 (by what path you handled failure), write the attempt limit as a number.