TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

Can You Go Back to Where You Were Yesterday?

Continue in TT Lab

In one line

Once you attach a checkpointer, the graph leaves the whole state at every step. That is why you can continue, rewind to the past, and branch from there. And for the same reason, the size of the state becomes the storage cost.

Why this was needed

Say an agent failed at the eighth step of ten. Without a checkpointer there is one way: start over from the beginning. If the first seven steps wrote files and sent mail, that happens again too.

The more common case is this. A user looks at the answer and asks, "what if a different source had been used at the third step?" If the state is not kept, there is no way to answer. The only way is to fix the input and run again from the start, and then even the results the first two steps made are created anew. You have to compare under the same conditions, but the conditions have changed.

A checkpointer solves both problems with one method. Every time a step ends, it stores one copy of the whole state at that point. Each stored point has a label (checkpoint_id), and if you take that label, you can go back to that spot at any time.

How it works

Using it takes two lines. You give a checkpointer when compiling, and give a thread_id when running.

from langgraph.checkpoint.memory import MemorySaver

app = graph.compile(checkpointer=MemorySaver())
config = {"configurable": {"thread_id": "user-42"}}
app.invoke({"topic": "가을"}, config)

thread_id is a name that points to one conversation. If you give the same value, it continues on top of the saved state, and if you give a different value, it starts fresh from nothing. The Persistence documentation calls this short-term memory — memory that continues only within one conversation.

What matters here is how many points are saved. When I ran a graph with three nodes chained in a line once in this lab environment (langgraph 0.2.60) and counted get_state_history(config), I got five. One at the point where the input was received, one each (three) as each node finished, and one at the point where everything was done. If one node is added, one checkpoint is added.

The list comes out newest first. If you want to read it in time order, you have to reverse it. A single snapshot contains the values at that point (values), the node that will run next (next), and the config that points to that spot.

Rewind — run again from the saved spot

If you use the config of a past snapshot as it is and give the input as None, it runs again from that spot.

snapshot = ...            # next 가 ("write",) 인 스냅샷
app.invoke(None, snapshot.config)

Giving the input as None is the key. If you give a new input, it starts fresh, and None means "continue with that saved state". Measured in practice, rewinding to before write added two checkpoints (write and review). Before plan it is three, before review it is one. It grows only by as much as it ran again.

Fork — change a value and go another way

If you change in one value at the same coordinates, a different branch is created from there.

forked = app.update_state(snapshot.config, {"angle": "비교"})
app.invoke(None, forked)

update_state puts one more checkpoint at that spot and returns a config with a new checkpoint_id. If you continue with that, a new branch is made.

There is one thing people often misunderstand here. The original branch is not erased. The checkpoints stay as they are, so you can read it at any time with the config of that point. But if you give only a thread_id and call get_state(config), what comes back is the end of the new branch. This is because the thread's "now" has moved. If you want to put the original result and the new result side by side and compare them, you must hold the config of that point in your hand before branching. Measured in this lab, after running once and then forking before write, the checkpoints went from five to eight — one at the spot where the branch was made, plus two nodes that ran again. Use time-travel explains these two as replay and fork.

What is saved is the whole state

There is no need to be confused about what goes into a checkpoint. It is the whole state. Not a part.

So if you put a big value in the state, the checkpoints grow along with it. In terms of time it varies with the machine and load, but the size is always the same for the same state, so you can just measure it. In this lab's graph I serialized the state and counted bytes. When nothing was put in, the sum of the five checkpoints was 399 bytes. When I put one 1,500-byte value in the state and ran the same graph, the sum became 6,447 bytes. The checkpoints actually holding that value were four of the five. A value put in once was stored four times.

This ratio grows as the number of nodes grows. With thirty nodes, a value put in once is stored nearly thirty times. So it is better to keep big original text or tables outside the state (in a file or a store) and put only a value that points to it in the state.

What it looks like in the field

First, the conversation does not continue. Either you did not attach a checkpointer, or you attached one and are creating a new thread_id every time. No error is raised. It just acts as if it were the first time, every time. Measured in practice, even running twice with the same config on a graph without a checkpointer started fresh both times.

Second, after branching, the original result is not visible. It was not erased; the thread's "now" moved. If you hold the config of the original point, it reads as it was.

Third, the storage grows noticeably. A graph that holds the original text whole in the state and runs many nodes ends up like this. To find the cause, you have to measure size, not time.

Fourth, when the Pod dies, everything disappears. As its name says, MemorySaver keeps things in memory inside the process. When a new process starts, nothing remains even if you ask with the same thread_id. It suits labs and tests but not production. If it has to last a long time, use a checkpointer that writes to an external store — which ones exist is written in Persistence and the graph API reference.

What really matters in practice

What you will do in the next lab

You grow /root/work/agckpt/ckpt.py one step at a time. First you confirm with MemorySaver and thread_id that the same conversation continues and different conversations are kept separately, and you count the checkpoints and read them as a list. Then you build a function that finds the coordinates to rewind to, rewind with those coordinates, change a value to make a branch, and read the original branch back. After that you serialize the state and count bytes to confirm in numbers how many places one big value rides in, and finally you reproduce with your own hands what disappears when you create a new ledger. The grader actually imports and runs your module, and pokes at it with different topics and values of different sizes each time.