TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

A Graph That Skips Approval Raises No Error

Continue in TT Lab

In one line

To stop in front of something irreversible, a breakpoint alone is not enough — without a checkpointer, the graph stops silently, and that case simply becomes something that never happened.

Why this was needed

When we launched a refund agent, we added "if it is over 100,000 won, a person looks at it before it goes out". It was a single line, interrupt_before=["settle"]. The demo went well too.

A week later, customer support got in touch. A large refund case did not appear in the approval-pending list, and the payout did not happen either. There is no error in the log. The agent had ended as a success.

The cause was one line. We passed only compile(interrupt_before=["settle"]) and did not attach a checkpointer. Measured directly in this lab image's langgraph 0.2.60, this is what happens.

체크포인터 없이 interrupt_before → 예외 없음. 결과는 {'amount': 100, 'trace': ['assess']}
                                   settle 을 지나지 않았고, 이어서 돌릴 방법도 없다

If an exception had been raised, it would have been caught that day. Because it quietly returned a half state, it took a week.

How it works

Pausing is built on top of saving. As the Persistence documentation says, the checkpointer leaves the state at every step. Pausing means "do not do the next step now; continue later from the state left behind". If there is nowhere to leave it, you cannot continue either.

So three things have to be present together.

After it pauses, you look at the state with get_state(config). .next holds "the nodes that will run next", and its being non-empty means "it is not finished yet".

state = app.get_state(config)
state.next            # ('settle',)  — 여기서 기다리는 중
state.values          # 그 시점의 상태 전체

When you continue, you give None in the input slot. app.invoke(None, config) means "there is no new input. Continue from where it was saved."

Values a person changes go through the reducer

What actually happens most often is that an approver reduces the amount or rejects it. The tool for that is update_state.

app.update_state(config, {"amount": 50000, "decision": "reject"})
app.invoke(None, config)

There are two places where you easily get caught here.

First, the values update_state puts in also go through that key's reducer. A key that overwrites is overwritten, but if you write to a key that appends (such as trace), it accumulates. So putting one line in trace to leave a mark of the human's edit is natural, and conversely, if you put something in "to replace the whole list", it grows against your intention.

Second, a fixed value can send the branch through again. update_state is recorded as "the node that ran last wrote that value". So if that node has a conditional edge attached, that condition is evaluated again. Measured directly in this lab, it shows up like this — if an approver reduces the amount of a 500,000 won case that is waiting for approval to 50,000 won, the case leaves the approval path and drops into the automatic processing path. Because needs_approval is called again and now says "approval not needed".

Whether this is a bug or a feature is decided by the business. If it is fine for it to go out as it is because it was reduced to a small amount, it is a feature, and if "a case that has come up in the approval-pending list must be seen through by a person", it is a bug. In the latter case, you have to separate the value used for judgment from the value used for execution — carry the originally requested amount separately, decide the branch by that, and use the corrected amount only for execution.

Stopping outside a node and stopping inside a node

interrupt() stops inside a node. When it stops, you can pass along a value to show to a person, and when you resume, the answer the person gave comes in as the return value of interrupt().

def confirm(state):
    answer = interrupt({"question": "이 환불을 승인합니까", "amount": state["amount"]})
    return {"decision": "approve" if answer == "yes" else "reject"}

app.invoke(Command(resume="yes"), config)

It is convenient, but there is a property you must know. When you resume, that node runs again from the beginning. Counting directly, it looks like this.

첫 실행 후   노드에 들어온 횟수 1
재개 후      노드에 들어온 횟수 2

So side effects placed before interrupt() (sending mail, external calls, incrementing a counter) happen twice. If you do not know this and put a payment before interrupt(), it is charged twice. The rule is simple — put only things that are fine to redo before interrupt(). Things that cannot be undone go after interrupt(), or in the next node altogether.

By contrast, interrupt_before stops before entering the node. So that node runs only once. Instead, there is no place to pick separately the value to show to a person, and you end up showing the whole state.

interrupt_before interrupt() inside a node
Where it stops Before the node Inside the node, at the point where it is called
Times that node runs 1 2 (on resume, again from the beginning)
Value handed to the person The whole state What you picked with interrupt(값) (the placeholder stands for the value)
Resume invoke(None, config) invoke(Command(resume=답), config) (the placeholder stands for the answer)

What it looks like in the field

First, leaving out the checkpointer. This is the incident above. No error is raised, so it lives long. Once you have built an approval path, always include a test that checks "does it really stop" with get_state().next.

Second, making everything need approval. If a person looks even at small, simple refunds, the queue backs up, and a backed-up queue ends up being looked at by nobody. Write the criterion in one place in the code (needs_approval), and make changing that criterion the same as changing the policy.

Third, having no way to express rejection. If there is only approval and no rejection, the approver rejects by "just not pressing". Then that case stays in the queue forever. Rejection must also be a result.

Fourth, having no approval record. Before who approved, you must record whether the requested value and the value actually sent out differ. If the approver changed the amount, that fact must be in the record so that you can explain it later.

What really matters in practice

What you will do in the next lab

You grow /root/work/aghitl/approve.py one step at a time. First you build the criterion and the branch that separate work that needs approval from work that does not, and you run a version that deliberately leaves out the checkpointer to see the silent stop with your own eyes. Then you attach the checkpointer and thread_id to make it really pause, continue, and continue again after the approver changes the amount or rejects. Finally you build a way to pause inside a node and count for yourself how many times that node is entered, and build a record that keeps the requested value and the executed value together.