AI Agents — A Graph, Not a Model
A Tool Is a Contract — Check Before the Call, Check Again After
In one line
Half of agent incidents come not from the model but from the side that calls tools. They call without measuring the arguments, and trust what comes back without measuring it.
Why this was needed
The way an agent with tools attached first breaks usually looks like this.
TypeError: stock_lookup() got an unexpected keyword argument 'sku_id'
The model made up one changed argument name, it was passed to the function as it was, and the exception popped out of the graph. This exception ends not a node but the whole run. The state, the path taken and why it happened are all gone.
The next way it breaks is worse because it is not an exception. The tool returned an empty list, and an answer was built on top of it. The user receives the sentence "stock in 3 places", and there is no error anywhere in the log.
The cause of the two incidents is the same. A tool call has no contract. If the code does not write down in one place which arguments it accepts with what type and range, and what it returns in what shape, the measuring code is scattered across nodes or absent altogether.
Where do you write the contract?
A contract must be not a document people read but a value a program reads. Tie the argument schema, the shape of what is returned and the actual function to one name and keep it as a single table.
TOOLS = {
"stock_lookup": {
"fn": stock_lookup,
"args": {"sku": {"type": "str", "required": True},
"region": {"type": "str", "required": True, "allowed": ["busan", "jeju", "seoul"]},
"limit": {"type": "int", "required": True, "min": 1, "max": 2}},
"returns": {"sku": {"type": "str"},
"rows": {"type": "list", "max_len": 2}},
},
}
If you do this, three things follow at once. First, you do not need code to measure arguments for each tool — one function that reads the table is enough. Second, you can pull out the tool descriptions to give the model from this table. Third, the door for calling tools narrows to one place. The tool-calling structure shown in Workflows and agents is also, in the end, a story about where you put this one narrow door.
Measure before calling
The most important thing in argument validation is the order: before calling. It is different from catching the exception after calling. The moment you call once with wrong arguments, money goes out, mail is sent, or an order is placed. A lookup tool can be undone, but a write tool cannot.
There are four things to measure: missing arguments, type, range and allowlist. If you add passing arguments that are not in the contract (typos, hallucinations), it is five.
There is one pitfall you always run into when measuring types in Python. bool is a subtype of int, so isinstance(True, int) is true. It means that even if True arrives in the count slot, it passes the "integer" check. So when measuring an integer, you must first filter out with isinstance(value, bool).
Let us also decide what validation returns. A single boolean is not enough. What is wrong and why must be left as a reason string so that it can be used next. If you put "which check" caught "which argument" in one line, like allowed:region, you can choose the response later from that one string.
The place that picks the arguments is originally the model
This reading and the next lab do not call a model. The node that picks the arguments is just written as rules, and in a real agent, that is the place where you ask the model — the model decides which tool to call and what to fill the arguments with. The rest of the structure does not change by a single line.
If anything, the more the side that picks is a model, the more this structure is needed. Rules are wrong in the same way each time, but a model is wrong differently each time. It swaps an argument name for a similar one, gives a number as a string, and makes up a region name that is not in the contract. So the value of this design is that two directions mesh in one place: pulling out the tool descriptions to give the model from the contract table, and measuring again against the same contract table the arguments the model returned.
Measure again after receiving
What a tool returns is a value someone else made. You must not treat it like the return value of our own function. There are four that you really meet often.
- The shape is different. A key is missing, or a string like
"4개"(the Korean for "4 items") arrives in the count slot. - It is an empty result. The list is empty, but it is not an error. If you pass this along as a success, you end up making up an answer.
- It is too big. You asked with
limit=2and twenty lines came. It is not rare for an old-version API to accept an argument and not honor it. - It is out of range. Negative stock, or 720 where it should be 14 days. 720 is usually a wrong unit — hours were put in the days slot. A unit error shows up like this as a type or range problem.
The function that measures these four is also built by reading the contract table. If you write result validation by hand inside the node, you miss it every time a tool is added.
Failure is a value, not an exception
Make the door for calling tools a function that does not let exceptions out.
{"ok": False, "value": None, "error": "args:allowed:region"}
If the shape returned is always the same, the node gets something to branch on. And if the reason written in error accumulates in the state, what to fix in the next attempt is decided in code. If the reason is args:allowed:region, change the region to the default, and if it is result:too_big, switch the tool from the old version to the new version. Conversely, if you throw it as an exception, all that remains is "something went wrong", so you cannot separate failures you can fix from failures you cannot.
There are clearly reasons that cannot be fixed, too. An empty result is one. Asking the same question again still gives an empty result. Retrying then only burns budget, and making up an answer is worse. Write the list of fixable reasons in the code, and give up right away on anything outside it. The attempt limit is a safeguard you put one layer above that.
Do not ask the same thing twice
An agent that loops keeps calling the same tool with the same arguments. Either the result did not remain in the state, or even if it did, the next node cannot find it. If the door for calling is in one place, putting a memory in front of that door is all it takes.
The key is 이름 + 정규화한 인자 (the name plus the normalized arguments). The same key has to come out even if the arguments were written in a different order, so you build it sorted, like json.dumps(args, sort_keys=True). And remember only successes. If you remember failures too, the path to fix and call again is blocked.
What it looks like in the field
First, the model makes up argument names. Arguments that are not in the contract arriving is not an accident but an everyday event. Block it, leave the reason, and ask again.
Second, the lookup succeeded but there is nothing inside. If you pass on looking only at ok, a sentence gets built on top of it. One line that classifies an empty result as a failure prevents this.
Third, an old-version API ignores the arguments. A response that does not honor the limit goes into the model's context as it is and eats tokens. Result validation is also cost management.
Fourth, the same lookup runs five or six times in one run. It shows up first on the bill. Either the result did not remain in the state, or even if it did, the next node could not find it.
Fifth, a single exception ends the whole run. If a KeyError raised inside a tool function goes out of the node and out of the graph, that case disappears entirely. The state accumulated up to the middle disappears with it, so when you run again you have to start from the beginning.
What really matters in practice
- Keep the tool contract in one place as a value. Name, arguments, the shape returned and the function in one table.
- Validate arguments before calling. What differs from catching after calling is whether you can undo it.
- Result validation is four things: shape, empty result, size and range. Units usually show up as a range problem.
- Return failure as a value and leave the reason in the state. That way the next attempt becomes a rule.
- Keep both the list of fixable reasons and the attempt limit.
What you will do in the next lab
You grow /root/work/agtool/tools.py one step at a time. You build the tool contract table, a function that measures arguments before calling and a function that measures the result after receiving it, and set up a call door that does not let exceptions out. Then you attach a memory so the same arguments are not called twice, build a graph that treats failure as a path, and finally build a version that reads the reason, fixes the arguments and tool, and tries again. The place that picks the arguments is replaced by rules, but a real agent asks the model in that place — the rest of the structure is the same. The grader actually imports your module, pokes at the validation functions with different values each time, and even counts how many times the tool body ran.