TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

Retry It, or Cap It? Deciding Both Up Front

Continue in TT Lab

In one line

Retry is switched on with one line, but a retry that does not decide what to redo is merely a device that repeats the same failure up to the limit. And that repetition eats the call budget as it is.

Why this was needed

Once you have built an agent as a graph, one day you end up writing a retrospective like this: "the payment gateway wobbled for about 30 seconds, and every case that came in during those 30 seconds ended in failure."

So you switch on retry. In LangGraph it takes one line per node.

graph.add_node("settle", settle, retry=RetryPolicy(max_attempts=3))

And next week you write the same retrospective again. It is one of two things.

These three things — what to redo, keeping the same request from being applied twice, and how much to spend — are each easy if learned separately, but if you do not bundle them, one of them is always missing.

The default retries less than you think

Writing out the arguments of RetryPolicy as they are gives this. They are the values I confirmed directly in this lab environment's langgraph 0.2.60.

RetryPolicy(initial_interval=0.5, backoff_factor=2.0, max_interval=128.0,
            max_attempts=3, jitter=True, retry_on=<기본 함수>)

What to notice is the last one, retry_on. If you open the body of the default function, it looks like this.

def default_retry_on(exc):
    if isinstance(exc, ConnectionError):
        return True
    if isinstance(exc, (ValueError, TypeError, ArithmeticError, ImportError,
                        LookupError, NameError, SyntaxError, RuntimeError,
                        ReferenceError, StopIteration, StopAsyncIteration, OSError)):
        return False
    ...
    return True

That is, the default policy does not redo ValueError. Neither RuntimeError nor OSError. What it does redo is ConnectionError and HTTP responses that ended in 5xx. Thinking about it, it is a reasonable default — a ValueError or TypeError is usually our code written wrongly, and such things give the same result even if redone a hundred times.

The problem is that exceptions we create usually inherit from ValueError or RuntimeError. If you make "the gateway is briefly unavailable" a class GatewayBusy(RuntimeError), then even with the policy attached, it never redoes anything. No error or warning appears. So if you move on trusting only the fact that you attached a policy, you write the same retrospective a few weeks later.

The fix is to write retry_on yourself. Give a tuple of exception classes, or give a function.

RetryPolicy(retry_on=(GatewayBusy,), max_attempts=3)
RetryPolicy(retry_on=lambda exc: isinstance(exc, GatewayBusy), max_attempts=3)

max_attempts is the total number of attempts, not the number of additional retries. If it is 3, that is the first try plus two more. And if it still fails after using all three, the last exception comes up as it is — no mark of having retried is attached to the exception. You have to count and leave that mark yourself.

Failures you may redo and failures you may not

This is the genuinely hard place. If you do not sort failures by kind, retry goes wrong in one of two ways — it redoes nothing, or it redoes everything.

One criterion is enough. If you send the same request again as it is, is there a chance of a different answer?

If you apply retry to a permanent failure, the cost triples and that is all. This is not something to take on faith in words; you have to look at it by count. If you count the attempt log, you can see as it is that one permanent failure piles up three lines of attempts. That number becomes your basis for explaining to others.

It must be idempotent before you redo

For retry to be safe, there is one more condition. Even if you send the same request twice, it must be applied only once.

A request that dropped by timeout is especially dangerous. You merely did not receive the response, and on the other side it may have already been processed. If you resend it as it is at that point, it is billed twice. So you attach a key to each request, and the receiving side uses that key to see "was this already done?" The key has to be decided by the side that creates the request — if the receiving side makes it fresh each time, there is no way to know whether two requests are the same request.

What you use as the key matters. It has to be a value that points to that one job, like an order number. If you use the time or a random value, it becomes a new key on each retry and idempotency disappears.

The budget counts retries too

Last is the budget. An agent calls far more than people expect. If you do not decide how many times to call tools to handle one case, one case uses up other cases' share too.

What is often left out here is the fact that a retry is also a call. If you say up to three attempts per case and run ten cases within one budget, then if the first two cases merely wobble, the later cases cannot even try. So it is more accurate to count the budget not in "how many cases" but in "how many times the tool body ran".

And what to do when the budget runs out remains. The answer is not retry. Give up, but leave a result. If you just send the exception up, the caller is left with only a stack trace, and nobody knows what was done and what was not. Exceeding the budget is a failure, but an expected failure, so it must be one line of the result record.

This lab uses RetryPolicy from the Types reference and the node settings from the Graph API overview. How to call tools and how to validate results are covered separately in the tools module, and here we look only at when it fails after calling. Fixing the arguments by reading the reason and calling again is something done inside the graph, and what is covered here is the framework-level retry that resends the same request as it is — the layers differ.

What it looks like in the field

First, you attached the policy but the retry does not happen. The exception is in the ValueError or RuntimeError family, so the default retry_on filtered it out. You can tell from the attempt log having only one line.

Second, it sends three times even for a card decline. You set retry_on too wide. You hit the gateway's usage limit first, so the cases that really should be redone get pushed back.

Third, the same case is billed twice. A request that dropped by timeout was resent, and the receiving side had no idempotency key. This incident is usually discovered first by accounting.

Fourth, the latter part of a batch did not run at all. The retries of the first few cases ate the whole budget. In the log, even "budget exceeded" is not visible — there is simply no record at all.

Fifth, covering it up by raising the limit. If you raise max_attempts to 10, you get past that day, but the cost of permanent failures becomes 10 times. What to raise is not the limit but the criterion for separating.

What really matters in practice

What you will do in the next lab

You grow /root/work/agbudget/budget.py one step at a time. First you reproduce that in a graph with no retry a failure is attempted only once, and confirm by attempt count that even if you attach a RetryPolicy, the default retry_on does not redo that failure. Then you state retry_on explicitly to make it redo, see by count that a permanent failure repeats up to the limit, and put in a separating function to bring that number back to 1. Next you use an idempotency key so the same request is not applied twice, count the call budget and give up but leave a result when it runs out, and finally confirm that when several cases share one budget, the retries of the earlier cases eat the share of the later ones. The grader actually imports your module, runs it with different keys, amounts and failure plans each time, and compares the attempt counts with values it computes separately.