TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

Streaming Sends Steps, Not Just Characters

Continue in TT Lab

In one line

What you stream out of an agent is not only tokens. Which node it is in now, what has newly been decided and the whole state after it ends each come out in a different mode, and the criteria for choosing differ.

Why this was needed

We attached an agent that tidies up meeting notes, and users said "it seems stuck". In fact it was running fine. It was a job that takes about 12 seconds, and the screen simply showed nothing during that time.

The first thought was "let's stream tokens one character at a time". But this agent calls the model only once, at the end. The moment there are tokens to stream is already the moment it is almost done.

What was needed was something else: streaming steps and intermediate results like "splitting the meeting notes into paragraphs now" and "found 4 decisions". And that is something the graph knows, not the model.

How it works

As the Streaming documentation sets out, one line of stream(입력, stream_mode=...) (the placeholder stands for the input) is enough. The shapes I measured directly in this lab image's langgraph 0.2.60 are these.

Mode Shape of one event How many come out
values The whole state (a dictionary) 1 input state + 1 per superstep
updates {노드이름: 그 노드가 돌려준 갱신} (node name: the update that node returned) 1 per node
debug {"type": "task"/"task_result", "step": 번호, "payload": {...}} (the placeholder stands for the number) 2 per node
custom Exactly what the node put in with writer(...) As many as the node called
If you give a list (모드이름, 값) tuples (mode name, value) Events of the chosen modes, mixed together

There are two places where it is easy to get confused.

First, updates is per node, not per superstep. Even if two nodes running side by side are in the same superstep, the events come out separately, one each. By contrast, values comes once after the superstep ends, so the results of the two side-by-side nodes come out after being merged.

Second, step in debug is the superstep number. Nodes running side by side have the same number. So if you want to know "which step is it now and what is running together in that step", you use this mode.

If you collect only the updates, can you rebuild the state?

You can. But you have to know the reducers.

for event in app.stream(입력, stream_mode="updates"):
    for node, update in event.items():
        for key, value in update.items():
            state[key] = value        # ← 이어 붙이는 열쇠에서 틀린다

If you overwrite an appending key like trace this way, only the last entry remains. The job the reducer did inside the graph, you have to do yourself when reconstructing outside.

But even if you imitate the reducer correctly, it still does not become completely identical. Measured directly in this lab image, it looks like this.

그래프에 노드를 split → keypoints → actions → compose 순으로 더함

updates 로 받아 이어 붙인 trace : ['split', 'keypoints', 'actions', 'compose']
values 의 마지막 이벤트의 trace : ['split', 'actions',   'keypoints', 'compose']

The positions of the two nodes running side by side are swapped. updates comes out in the order the nodes were added to the graph, while the reducer merges the values of the same superstep in node-name order. Both always give the same answer (they do not wobble when repeated), but they differ from each other.

So the rule is this — if the data is such that order carries meaning, do not use what you reconstructed from updates as the final version. It is fine for drawing the screen, but for the final state to save or compare, use the last event of values (or the answer of invoke).

Setting that aside, the practical choice usually splits like this — what the screen must draw cumulatively (a progress log, a partial list) is received as updates and merged by the front end, and what just needs the whole final state refreshed uses the last event of values. Since values sends the whole state every time, the amount transferred is large — if the state contains a big value, that value goes out again at every superstep.

Pieces a node sends out directly

There are things you do not want to leave in the state but do want to show people. A progress indicator like "processing paragraph 3". If you put it in the state, it rides in the checkpoint and comes along when you resume.

The custom mode is the place for that. A node takes a writer argument and sends it out directly.

def compose(state, writer: StreamWriter):
    writer({"stage": "compose", "points": len(state["points"])})
    return {"summary": ...}

What you put in with writer(...) does not remain in the state and goes only to the listening side. So it suits progress indicators, partial text and debug hints.

What you streamed and the answer of invoke are the same

The last event of values is the same as what invoke() returns. You can confirm it by comparing directly, and that they are the same matters — it means there is no reason to run twice, such as drawing the screen by streaming and recording separately with invoke.

What it looks like in the field

First, trying to stream at a moment when there is nothing to stream. In a graph that calls the model once at the end, if you attach only token streaming, the silence before it stays as it is. You have to stream the steps.

Second, sending a big state out with values at every step. If the state contains the original text, that original text goes out again at every superstep. It is better to send only what the screen needs with updates or custom.

Third, merging updates by overwriting. This is the reconstruction trap above. Only one footprint remains.

Fourth, putting progress indicators in the state. The checkpoints grow, and when you resume, old progress indicators come back to life.

What really matters in practice

What you will do in the next lab

You grow /root/work/agstream/stream.py one step at a time. You stream the same graph in five ways, values, updates, debug, several modes together and custom, and count the number and shape of the events yourself. You gather only the updates to rebuild the last state and confirm what goes wrong with an appending key. Finally you compare whether the last state you streamed and the answer of invoke are the same, and organize what to show and when, leaving it as a record.