AI Agents — A Graph, Not a Model
An Agent Is a State Machine, Not a Model
In one line
An agent is not a model but a state machine. The hard part is not the place where you call the model, but the places where you decide when to stop, what to carry along, and where to go when something fails.
Why this is needed — where a demo and a product part ways
Agent demos usually run fine. You put in one question, it calls a tool, and an answer comes out. But the moment it becomes a product, three things blow up at once.
- It does not stop. If a tool cannot give the answer it wants, it keeps retrying. The token bill tells you so.
- It makes up what it does not know. The lookup came back empty-handed, yet it produces an answer.
- You do not know why it did that. The same question gets a different answer, and there is no record of which path it took.
None of the three is a model problem. They are graph problems. That is why this course does not call a model and deals only with the graph. Attaching a model is a matter of swapping one node, and that part is not hard.
Is state overwritten, or does it accumulate?
In LangGraph, a node returns a dictionary containing only the part it changes. How to merge it is decided by the state type.
class State(TypedDict):
question: str
steps: Annotated[list, operator.add] # 쌓인다
tries: int # 덮어쓴다
Without Annotated[list, operator.add], only the value returned by the last node remains. Then the "path taken" disappears, and you cannot ask later what happened. What to accumulate and what to overwrite is a design decision.
What becomes testable once you move to a graph
If you write an agent as one big loop, there is no place to ask what went wrong. The real reason for splitting into a graph is that each node can be tested on its own.
Make nodes close to pure functions. If a node takes the state and returns part of the state, you can detach just that node, give it inputs and look at the result. Keep the model call inside the node but make it injectable, so that in tests you swap in a fake that returns fixed answers.
def plan(state: State, llm=None) -> dict:
llm = llm or default_llm
out = llm.invoke(state["messages"])
return {"plan": parse_plan(out), "step": state["step"] + 1}
Put the branching in a condition function, not in a node. If you keep the function that decides where to go next separate, that function alone lets you build a table of "in this state, call the tool; in that state, end" and test it. You can check the whole flow without calling a model.
State the rule for what accumulates and what is overwritten. Conversation history accumulates, and the current plan is overwritten. If you write this in the type, you will not get confused as you add nodes. If you overwrite something that should accumulate, the context disappears, and if you accumulate something that should be overwritten, the prompt swells.
Saving intermediate state brings resumption and human intervention together. If you save the state at every node, you can stop midway, ask a person, and continue from their answer. Work that needs approval (sending email, payment, deletion) is cut off at this point.
Leave observations per node. Only if you keep how many seconds each node took and how many tokens it used do you know where the slow and expensive places are. If you measure only the total elapsed time, you cannot find where to improve.
In the field
When an interview asks "have you built an agent?", what they actually want to hear is not a framework name but these three things: what you set as the ending condition, where you send a failed tool call, and what you left on record.
And this is not a story only about agents. Retry limits, failure paths and observation records are judgments you have always made in backends. LangGraph just lets you write those judgments down as a graph.