AI Agents — A Graph, Not a Model
Nodes Do the Work, State Remembers It
In one line
In an agent, the most expensive decision to fix later is the state design. And half of the state is not the list of keys but the merge rules (reducers).
Why this was needed
Once you build an agent as a graph, the first ill-mannered bug you meet is this: you passed through several nodes, but only the last entry of the path taken remains.
class State(TypedDict):
trace: list # 리듀서가 없다
def intake(state): return {"trace": ["intake"]}
def lookup(state): return {"trace": ["lookup"]}
def finish(state): return {"trace": ["finish"]}
After passing through the three nodes in order, if you open trace it is ["finish"]. It is not ["intake", "lookup", "finish"].
No error is raised. So this bug eats half a day as you suspect the logging code, wondering "why isn't the log being kept". The real cause is not the logging but that that key has no merge rule.
How it works
LangGraph's state is a structure in which each key has one channel. A node returns not the whole state but a dictionary containing only the keys it wants to change, and the graph applies it channel by channel. Keys it did not touch stay as they were — so even if a node returns only {"answer": ...}, question does not disappear.
The way of applying is the reducer. The rule set out in the Graph API overview is two lines.
- If you write no reducer, it overwrites. The new value replaces the old one.
- If you write one, it merges as that function decides.
Annotated[list, operator.add]appends, and a function you wrote yourself merges as that function decides.
A reducer is an ordinary function that takes two arguments, (지금까지의 값, 노드가 돌려준 값) (the value so far and the value the node returned), and returns the new value. There is nothing special about it. So you can collect the "merge rules" in one place in the code, and that place is the state definition.
def merge_sources(old, new):
out = list(old or [])
for item in new or []:
if item not in out:
out.append(item)
return out
class State(TypedDict, total=False):
trace: Annotated[list, operator.add] # 이어 붙인다
sources: Annotated[list, merge_sources] # 중복 없이 이어 붙인다
used_calls: Annotated[int, operator.add] # 더한다
Here is one pitfall confirmed by measurement. You cannot pass a built-in function directly as a reducer. If you write Annotated[int, max], it dies when the graph is compiled with ValueError: no signature found for builtin max. LangGraph inspects the reducer's signature, and Python built-in functions have no signature to inspect. Wrap it one level, like def keep_max(old, new): return new if new > old else old.
State grows — that is the cost
Once you attach reducers, the problem on the other side arrives. trace, sources and messages all grow without end.
Growing itself is fine, but the problem is that the checkpointer saves the state in full. If a node runs 30 times during one run, 30 checkpoints are created, and each of them holds the entire state at that point. If the state becomes ten times bigger, the storage also becomes ten times bigger.
So for a key that grows, write the limit inside the reducer.
def keep_recent(old, new):
return (list(old or []) + list(new or []))[-3:]
If you do it this way, nodes can just write, and the limit is enforced in one place only. It is better than scattering "trim if too long" across every node — if you scatter it, missing a single spot makes it leak silently.
Decide the entrance and the exit separately
The state gets mixed with things that must not be visible from outside: the original text the user uploaded, intermediate notes, internal scores. If you give input and output schemas separately with StateGraph(State, input=InputState, output=OutputState), nodes see the whole state while the outside receives only the keys of the output schema.
If you do not do this, you end up deleting "why does the response carry internal notes" on the screen side. The deleting code falls behind whenever a new key is added, and in the end it leaks. It is always safer to decide what goes out as an allowlist.
What it looks like in the field
First, the path taken remains as a single entry. It is exactly what you saw above. The reducer was left out, and the symptom appears as "the log isn't kept".
Second, the same source is written twenty times. If an agent that loops keeps referring to the same document, sources is filled with the same value. If you use a reducer that removes duplicates instead of operator.add, it ends right there.
Third, the budget is scattered across nodes. If the code that counts calls lives in three nodes, fixing one leaves the others behind. If you set used_calls: Annotated[int, operator.add] and have nodes return only {"used_calls": 1}, the counting rule stays in one place.
Fourth, checkpoints grow and resuming gets slow. A graph that holds the original text whole in the state and runs dozens of times ends up like this. It is better to keep the original outside (a file or a store) and put only a value that points to it in the state.
What really matters in practice
- Decide for each key "overwrite or merge", and write that answer in the state definition. If you do not write it, overwriting becomes the answer — what matters is whether that is intended or not.
- For a key that grows, put the limit inside the reducer. Not in the node, but in the reducer.
- Decide what goes outside as an allowlist. The output schema is that place.
- State rides in checkpoints. Before putting something in the state, ask "is it okay to save this at every step?"
What you will do in the next lab
You grow /root/work/agstate/state.py one step at a time. First you reproduce with your own hands a key with no reducer being overwritten, then you attach, in turn, an appending reducer, a duplicate-removing reducer, a reducer that keeps only the maximum, and a reducer that trims a window. At the end you split the input and output schemas to narrow the keys that go outside, and leave a record of why you decided that way. The grader actually imports and runs your module, and pokes at the reducers directly with different names and numbers each time.