Cinematic Short Agents in Production 52 sec Failure mode + fix

The Runaway Agent

Your agent looped all night. $10k gone.

Why it runs away

Nothing stops the loop, and each retry resends a bigger context. Cost grows faster than the step count.

The fix

Cap steps and tokens per run, kill repeated tool calls, back off failing tools, and alert on spend.

Context

A developer kicks off an autonomous coding agent before bed ("fix the build overnight"). The agent runs a ReAct-style loop: call the LLM, pick a tool, run it, append the result to the conversation, repeat until the model says it is done. It is billed per input and output token by an LLM API, and nobody is watching until morning.

Architecture

ComponentRoleNotes
DeveloperStarts the runAsynchronous: no human in the loop overnight
AI agent (loop runner)Calls the LLM, runs tools, appends resultsThe villain here: no max steps, no budget, no loop detection
Tool (e.g. a CI or test-runner API)Does the workStarts failing (outage, bad credentials, rate limit), returning errors
Context (conversation history)Everything the model has seen so farEvery tool call and error is appended; the whole history is resent each turn
LLM APIGenerates the next actionBilled per token; input tokens dominate as the history grows
LLM gateway (the fix)Proxy in front of the LLM APIWhere per-run and per-key budgets, rate limits and spend alerts live

Request flow

  1. The developer submits the task; the agent sends system prompt + task to the LLM.
  2. The LLM picks a tool call; the agent runs it and appends the result to the history.
  3. The agent sends the entire history back to the LLM for the next step.
  4. Repeat until the model returns a final answer (or something stops the loop).

The failure mode

Around 2 AM the tool starts failing: every call returns the same error. The model reads the error and decides the reasonable thing is to try again, perhaps with a trivial variation. Nothing in the loop stops it:

  • No step cap: the loop only ends when the model declares success, which never happens.
  • No budget: nobody tracks tokens or dollars per run.
  • No loop detection: the same tool is called with the same arguments hundreds of times.
  • No backoff: retries fire back to back, as fast as the LLM responds.

Each retry appends the call and the (often long) error text to the history, and each new LLM call resends all of it. If each step adds roughly k tokens, step n costs about n·k input tokens, so N steps cost about k·N²/2 in total: doubling the run length roughly quadruples the bill. By morning the run has made thousands of calls and burned a five-figure sum. (The $10k figure is illustrative; the exact number depends on model price, error size and how long the loop runs. Prompt caching lowers the price of the repeated prefix but does not stop the growth.)

It is easy to miss because each single call looks normal and cheap, dashboards usually aggregate per day, and the agent's log looks like diligent work ("retrying…").

Why it happens

The loop's only exit condition is the model's own judgment, and a failing tool gives the model a reason to keep trying. Without external limits, cost is unbounded, and history resending makes it grow superlinearly.

The fix

Layered guardrails, so no single missing check lets a run run away:

  1. Max steps per run. A hard iteration or turn cap in the agent runner. Frameworks ship one: LangGraph's recursion_limit (default 25) and the OpenAI Agents SDK's max_turns (default 10). Don't raise it to "unlimited".
  2. Per-run token and dollar budget with a hard kill. Count tokens on every call; stop the run (and save state for a human) when it crosses the budget.
  3. Loop and repeat detection. Hash (tool name + arguments); if the same call repeats N times, or the same error comes back N times, stop or escalate instead of retrying.
  4. Exponential backoff with a circuit breaker on tools. After a few consecutive failures, open the breaker and fail the tool fast for a cool-down period, so the agent gets a clear "tool unavailable, stop" rather than another error to chew on.
  5. Spend limits and alerts at the LLM gateway. Per-key or per-project budgets, rate limits and an alert when spend per hour spikes. This is the backstop that catches bugs in every agent, not just this one. Also: trim or summarise old tool errors so the context doesn't grow without bound.

Trade-offs

  • Step caps and budgets can cut off legitimately long tasks; make them per task type and resumable, not global.
  • Loop detection has false positives (polling a job status is a legitimate repeat); key it on arguments and results.
  • Circuit breakers need tuning (thresholds, cool-down) and add state to the runner.
  • Gateway budgets add a hop and one more service to run, but centralise cost control.

Numbers worth knowing

  • LangGraph default recursion_limit: 25 steps. OpenAI Agents SDK default max_turns: 10 (framework docs).
  • Cost of a loop that resends history grows roughly with the square of the step count (arithmetic, see above).
  • "$10k overnight" is illustrative, not a specific incident.

Coming next in the series: LLM Gateways

Found this useful?

Subscribe for the next episode, or share it with the person who owns this part of your stack.

Keep going