TT Lab
Get started
Learn Learning paths Courses

AI Agents — A Graph, Not a Model

We Retried the Declined Card Three Times

Continue in TT Lab

Goal

You attach a RetryPolicy to a node and make it redo only the failures that may be redone. You attach an idempotency key so the same request is not applied twice, count the call budget, and when it runs out, give up instead of retrying but leave a result. All judgments are by count — time and speed are not measured.

Why it matters

Retry is switched on with one line. So people switch it on and do not check. But the default retry_on of RetryPolicy does not redo exceptions like ValueError, TypeError, RuntimeError and OSError (this is a fact confirmed in this lab environment's langgraph 0.2.60). The exceptions we create usually inherit from that family, so even with the policy attached, only one line remains in the attempt log. No error or warning appears. Conversely, if you set retry_on wide, even permanent failures repeat up to the limit. One card decline piles up three lines of attempts, and that much eats the share normal cases would use. So you need a criterion for separating "failures that may be redone" — if you send the same request again as it is, is there a chance of a different answer? For retry to be safe, there is one more condition. A request that dropped by timeout may merely not have gotten a response and may already have been processed on the other side. If you resend it as it is, it is billed twice. You must attach a key to each request and have the receiving side use that key to see "was this already done?" Finally, retries eat the call budget as they are. If the earlier cases merely wobble, the later cases cannot even try. You count the budget not by the number of cases but by the number of times the tool body ran, and when it runs out, you leave it as one line of the result record instead of raising an exception. The grader does not trust the explanations you wrote. It actually imports your module, runs it with different keys, amounts and failure plans each time, and compares the number of lines piled up in ATTEMPTS and the amounts applied to the ledger with values the grader computes separately.

Steps

  1. In /root/work/agbudget/budget.py, create MAX_ATTEMPTS, TransientError, PermanentError, ATTEMPTS, LEDGER, PLAN, reset(), applied_total(), charge(), State, build_plain() and run_once(). There is no retry, so if it fails, there is only one attempt.
  2. Add DEFAULT_RETRY, build_default_retry() and run_default(). You attached a RetryPolicy without retry_on to the node and confirm that there is still only one attempt.
  3. Add RETRY_ALL, build_retry_all() and run_retry_all(). If you state retry_on=(ValueError,) explicitly, transient failures come back to life, and if you exceed the limit, the last exception comes up as it is.
  4. Add is_retryable(exc), RETRY_SPLIT, build_split() and run_split() so that permanent failures are not redone. The attempt count of a permanent failure must drop from MAX_ATTEMPTS to 1.
  5. Put an idempotency key into charge(). Even if it is called twice with the same request_key, only one line remains in the ledger, and the second call returns that same receipt with the duplicate mark.
  6. Add BudgetExhausted, call_with_budget(), build_budgeted() and run_budgeted(). If no budget remains, do not call the tool, and even when giving up, leave a record whose outcome is "예산초과" (the Korean word for "budget exceeded").
  7. Add settle_batch(orders, budget) so that several cases share one budget. Cases after the budget runs out are left as skipped.
  8. Record the counts you measured in /root/work/agbudget/budget_report.json and /root/work/agbudget/budget_report.md.

Notes

Without retry, there is only one attempt

In /root/work/agbudget/budget.py, create MAX_ATTEMPTS, TransientError, PermanentError, ATTEMPTS, LEDGER, PLAN, reset(), applied_total(), charge(), State, build_plain() and run_once(). build_plain() returns a compile()d graph with no retry attached.

Each time charge is called, it first appends one line to ATTEMPTS and then looks at PLAN — because failed attempts must be counted too. The node must not swallow the failure and must let the exception out as it is. If you catch it with try/except inside the node, retry never kicks in. run_once wraps invoke in try/except and turns the exception into a record. If a node name and a state key overlap, compiling dies, so keep them different.

Why does it try only once even though I attached a policy

Add DEFAULT_RETRY = RetryPolicy(max_attempts=MAX_ATTEMPTS, initial_interval=0.01, backoff_factor=1.0, jitter=False) and build_default_retry() and run_default(). Do not write retry_on. Confirm that even when you run one transient failure, there is still only 1 attempt.

Use from langgraph.types import RetryPolicy and attach it with add_node(..., retry=DEFAULT_RETRY). This step deliberately reproduces something that does not work — the default retry_on does not redo the ValueError family. The grader looks both at whether the policy is actually attached to the node (the nodes["settle"].retry_policy of the graph after compile()) and at the attempt count. So you must not just return build_plain() as it is.

Write down what to redo yourself

Add RETRY_ALL = RetryPolicy(retry_on=(ValueError,), max_attempts=MAX_ATTEMPTS, ...) and build_retry_all() and run_retry_all(). If it succeeds after n transient failures, the attempts are n+1, and if it exceeds MAX_ATTEMPTS, the last exception comes up as it is.

Give retry_on a tuple of exception classes — do not forget the comma even when giving just one class. max_attempts is the total number of attempts, so 3 means the first try plus two more. If it still fails after using up the limit, the last exception comes out as it is, with no mark of having retried attached. You have to leave that mark yourself by counting ATTEMPTS.

Do not send a permanent failure three times

Create is_retryable(exc) and add RETRY_SPLIT = RetryPolicy(retry_on=is_retryable, ...), build_split() and run_split(). The attempt count of a permanent failure must drop from MAX_ATTEMPTS to 1.

You can give retry_on a function (예외) -> bool as well (the placeholder stands for the exception). There is one criterion for separating — if you send the same request again as it is, is there a chance of a different answer? Make only the kinds that may be redone true, and make unknown exceptions false. If you make true the default, a permanent failure silently leaks through each time it arrives under a new name. The grader runs the same permanent failure through run_retry_all and run_split and compares the attempt counts.

Even if it arrives twice, apply it once

Make charge(), when it receives the same request_key twice, not create a new line in the ledger and return that same receipt with duplicate true. The number of ledger lines and applied_total() must not double.

First scan the ledger by the key to see whether it already exists. seq must stay the same too — if you assign a new number, the same case looks like two cases. You must look it up after the place where failures are raised. What matters is that the key is a value the calling side decided. The grader mixes in two different cases and a duplicate delivery of one of them, and looks at the number of ledger lines, the total and seq together.

Give up when it runs out, but leave a record

Add BudgetExhausted, call_with_budget(request_key, amount, budget), build_budgeted() and run_budgeted(request_key, amount, budget). If len(ATTEMPTS) is at least budget, do not call the tool, and in that case return a record whose outcome is "예산초과" (the Korean word for "budget exceeded").

The budget is not the number of cases but the number of times the tool body ran. Make sure is_retryable does not see BudgetExhausted as true — redoing a failure caused by having no budget is pointless. The budget check has to come before calling charge so that the attempts do not grow by even one slot. run_budgeted must not let the exception out and must turn it into a record holding outcome.

The retries of the earlier cases eat the share of the later ones

Add settle_batch(orders, budget). Several cases share one budget, and when no budget remains, the remaining cases do not call the tool and go into skipped. The returned value is {"done", "failed", "skipped", "calls", "applied", "budget"}.

You must not call reset() for each case — the budget is shared by the whole batch, so ATTEMPTS has to continue. Look at the remaining budget once before starting a case too, and if there is none, do not call run_budgeted at all. done, failed and skipped hold only the request_key values, in order. The grader computes the same rule separately and compares the lists and calls.

Record the counts you measured

Write max_attempts, no_retry_attempts, default_policy_attempts, retry_all_permanent_attempts, split_permanent_attempts, split_transient_attempts, duplicate_delivery_entries, batch_budget, batch_calls, batch_done, batch_skipped and batch_applied in /root/work/agbudget/budget_report.json, and write /root/work/agbudget/budget_report.md in four sections: ## 무엇을 다시 해도 되는가 (what may be redone), ## 상한을 어디에 두었나 (where you put the limit), ## 같은 요청이 두 번 와도 안전한 이유 (why it is safe even if the same request arrives twice) and ## 예산이 바닥났을 때 무엇을 남기는가 (what you leave when the budget runs out).

Do not write the numbers by hand; fill them in with values you get by actually running your module. split_transient_attempts is the attempt count when it succeeded after 2 transient failures, and duplicate_delivery_entries is the number of lines left in the ledger after sending twice with the same key. Run the batch with exactly the orders, plan and budget set in the notes section of the instructions. The grader computes the same things separately and compares.