TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

It Was Saved, So You Can Go Back to It

Continue in TT Lab

Goal

In a graph with a checkpointer attached, you count directly what is left where. You confirm that the same thread_id continues, read the checkpoints as a list, rewind to a past coordinate, change a value to make a branch, and serialize the state to count bytes. Finally you reproduce what disappears when you create a new ledger.

Why it matters

The checkpointer leaves the whole state at that point at every step. This one sentence explains both the good thing and the bad thing. The good side is that you can continue. Even if it fails at the eighth step you do not need to start over, and you can answer the question "what if a different source had been used at the third step?" under the same conditions. If you invoke(None, ...) with the config of a past snapshot, it runs again from that spot, and if you change a value with update_state, a branch forms there. The original branch is not erased at that point, but the thread's "now" moves to the new branch, so to see the original result again you must hold the config of that point in your hand. The bad side is that what is saved is the whole state. If you put one big value in the state, that value is stored repeatedly as many times as there are checkpoints. In this lab you measure that not in time but in bytes — time varies with the machine and load, but size is always the same for the same state. The grader does not trust the explanations you wrote. It actually imports your module, runs the graph, counts the checkpoints itself, and compares with the values you returned. The topic and the size of the value you put in change with each run.

Steps

  1. In /root/work/agckpt/ckpt.py, create ANGLES, State, three nodes (plan, write, review), build_graph(checkpointer=None), new_saver() and cfg(thread_id). The same thread_id must continue and a different one must be kept separately.
  2. Add history(app, config), history_len(app, config) and pending(app, config) to read and count the checkpoints from the oldest.
  3. Add snapshot_before(app, config, node) to find the most recent checkpoint at which that node has not yet run.
  4. Add replay(app, config, node) to run again from that coordinate with invoke(None, snapshot.config).
  5. Add fork(app, config, node, values) to make a branch with update_state and also return the end of the original branch.
  6. Add state_bytes(values), history_bytes(app, config), weigh(topic, payload) and the state key bulk to measure the size of the checkpoints in bytes.
  7. Add lost_demo(topic) to reproduce that when you create a new ledger, nothing remains even with the same thread_id.
  8. Record the values you measured in /root/work/agckpt/ckpt_report.json and /root/work/agckpt/ckpt_report.md.

Notes

The same thread continues and a different thread is kept separately

In /root/work/agckpt/ckpt.py, create ANGLES, State, three nodes (plan, write, review), build_graph(checkpointer=None), new_saver() and cfg(thread_id). A graph given a checkpointer must stack on top of the earlier notes when run again with the same thread_id, and a different thread_id must start from empty.

One line, graph.compile(checkpointer=checkpointer), is all there is to it. If you make cfg(thread_id) a small function that returns {"configurable": {"thread_id": thread_id}}, the later steps are easier. You have to attach operator.add to notes for the continuation to be visible — if you do not, it is saved but the values are overwritten, so you cannot tell it continued.

A checkpoint is left at every step

Add history(app, config), history_len(app, config) and pending(app, config). history() lays out the checkpoints from the oldest, and pending() returns a list holding, for each snapshot, the name of the node that will run next (an empty string at the spot where everything is done).

app.get_state_history(config) is a generator and comes out newest first. Take it with list() and apply reversed(). A snapshot's next is a tuple — if it is empty, there is nothing more to run. When you count, it is two more than the number of nodes, because there is one more at the point where the input was received and one at the point where everything was done.

Find the coordinates to rewind to

Add snapshot_before(app, config, node). It returns the most recent checkpoint at which that node has not yet run, and None if there is no such spot.

The condition to look for is that the snapshot's next is (node,). If you scan the list laid out from oldest to newest from the back, you meet the 'most recent' first. What you return is the snapshot itself — what the later steps use is the config inside it, and that contains a checkpoint_id in addition to the thread_id.

Run again from that spot

Add replay(app, config, node). Find the coordinate with snapshot_before, continue with app.invoke(None, snapshot.config), and return {"from": node, "result": 결과, "added": 늘어난 체크포인트 수} (the placeholders stand for the result and the number of checkpoints added). If there is no coordinate, raise ValueError.

Giving the input as None is the key. If you give a state, it starts fresh, and None means 'continue from that saved point'. added is the value you get by measuring history_len before and after the rewind and subtracting — it grows by as many nodes as ran again.

Change a value and go down another branch

Add fork(app, config, node, values). Grab the end of the thread before the fork first, make a branch with app.update_state(snapshot.config, values), and then continue with that config. Return {"forked_config": ..., "result": ..., "original": 분기 전 끝의 값 딕셔너리, "total": 지금 체크포인트 수} (the placeholders stand for the dictionary of values at the end before the fork and the current number of checkpoints).

update_state returns a config with a new checkpoint_id. If you throw it away and continue with the original config, no fork happens. And after forking, the 'now' you get by asking with only the thread_id points to the end of the new branch, so you have to read the original result separately with the config of the snapshot you grabbed before the fork.

Measure by size

Add the state key bulk and state_bytes(values), history_bytes(app, config) and weigh(topic, payload). state_bytes is the UTF-8 length of the string from json.dumps(values, ensure_ascii=False, sort_keys=True, default=str). weigh runs the same graph once without bulk and once with payload in bulk, and returns {"checkpoints": ..., "thin_total": ..., "fat_total": ..., "payload_bytes": ..., "carrying": ...}.

carrying is the number of checkpoints whose size difference is at least payload_bytes when the two lists are compared spot by spot — meaning the number of spots that actually hold that big value. When you run twice, use separate ledgers (call new_saver() twice). Do not measure time — it varies with the machine and load. Bytes are always the same for the same state.

If you create a new ledger, it disappears

Add lost_demo(topic). Build a graph with a new ledger, run the run-1 thread once and measure the number of checkpoints, then build one more graph with yet another new ledger and ask with the same thread_id. Return {"kept": ..., "after_restart": ..., "values_after_restart": ...}.

MemorySaver keeps things in memory inside the process. Creating a new MemorySaver is the most honest way to imitate the process starting again. If you ask a new ledger, no error is raised and an empty value comes back — that is why this problem is quiet. That is the reason it suits labs and tests but not production.

Record the values you measured

Write topic, checkpoints_per_run, pending, base_score, fork_angle, fork_score, checkpoints_after_fork, payload_bytes, carrying and survives_restart in /root/work/agckpt/ckpt_report.json, and write /root/work/agckpt/ckpt_report.md in four sections: ## 무엇이 저장되는가 (what is saved), ## 되감기와 분기는 어떻게 다른가 (how rewind and fork differ), ## 크기로 잰 것 (what you measured by size) and ## 사라지는 것 (what disappears).

Do not write the numbers by hand; fill them in with values you get by actually running your module. topic is the topic string you used exactly as it is, and fork_angle is the name of an angle other than 기본 (the Korean word for "basic") — the grader recomputes the score from those two and compares. Fork before write. payload_bytes must be at least 300 for the difference to be visible. survives_restart is true or false.