AEGIS OSBlog
OCT 05, 2026

Circuit Breakers for AI Agent Loops: Preventing the Infinite Spend Bug

By Quinn · 5 min read

In the early days of distributed systems, we learned that a single failing service could take down an entire cluster through a cascade of retries. We solved this with the circuit breaker pattern. Today, we are facing a new version of the same failure mode: the infinite spend bug.

An AI agent in production is not just a script; it is a loop. When that loop loses its way, it does not just crash. It spends.

The Anatomy of a Runaway Loop

The infinite spend bug manifests in three primary ways in production environments.

First is the retry storm. An agent attempts to use a tool, the tool returns a transient error, and the agent decides the best course of action is to try again. Without a backoff strategy or a hard cap, the agent can generate hundreds of requests in seconds, each one costing tokens for the prompt and the failed completion.

Second is recursive spawning. In multi-agent architectures, a supervisor agent might delegate a task to a sub-agent. If the sub-agent fails to provide a satisfactory answer, the supervisor might spawn another sub-agent to check the first one, leading to a tree of recursive calls that grows until the infrastructure or the credit card gives way.

Third is the reasoning loop. This is the most subtle. The agent gets stuck in a "thought" cycle where it repeatedly analyzes the same data without ever deciding to call a tool or return a final answer. It is the LLM equivalent of a while(true) loop, but every iteration costs five cents.

Mapping the Pattern to Agents

Michael Nygard and Martin Fowler popularized the circuit breaker to prevent these cascades. The pattern uses three states:

  1. ·Closed: The agent operates normally. The breaker monitors metrics (tokens, steps, time).
  2. ·Open: A threshold is crossed. The breaker trips, immediately failing all further calls to the agent without hitting the LLM.
  3. ·Half-Open: After a timeout, the breaker allows a single "test" iteration. If it succeeds, the breaker closes; if it fails, it opens again.

For agents, we do not just monitor error rates. We monitor resource consumption.

Three Essential Breakers

Every production agent stack needs three specific types of cutoffs.

1. Token Budget Limits

This is your hard FinOps line. You set a maximum dollar amount or token count for a single session. If an agentic run for a single user request hits $2.00, the breaker trips. This prevents a single bug from turning into a five-figure surprise on your monthly bill.

2. Iteration and Step Caps

Even if tokens are cheap, latency is not. An agent that takes 50 steps to solve a problem that should take 5 is a broken agent. Iteration caps ensure that the agent converges on a solution within a reasonable number of turns.

3. Wall-Clock Time Cutoffs

LLMs are non-deterministic. Sometimes a reasoning chain just hangs or takes an eternity to stream. A wall-clock timer ensures that the user (or the calling system) gets a "timeout" error rather than waiting indefinitely for an agent that has lost the plot.

Placement Matters: The Orchestrator Layer

A common mistake is putting the loop logic inside the agent's prompt or its internal code. If the agent's reasoning is what is failing, you cannot trust the agent to police itself.

The circuit breaker must live in the orchestrator layer. It should be a wrapper around the agent's execution loop that has the authority to kill the process regardless of what the LLM "wants" to do next. This separation of concerns is what makes the system composable and resilient.

Graceful Degradation

When a breaker trips, the system should not just throw a 500 error. You need a strategy for what happens next.

In a well-architected system, a tripped breaker moves the task to a dead-letter queue. This allows for human escalation or a fallback to a simpler, non-agentic heuristic. It preserves the state of the runaway loop so your engineering team can audit exactly why the agent started spinning.

The Cost of Neglect

Consider the math of a runaway incident. If you have 10 agents running in parallel, and each gets stuck in a loop making 1,000 tool calls at an average cost of $0.01 per call, you are losing $100 per incident. Without breakers, that incident can run for hours.

Production engineering is about managing the tail risks. In the world of AI agents, the tail risk is an infinite loop with a direct line to your bank account.

Implementation Example

Here is a minimal pattern for an iteration-based circuit breaker in a TypeScript-based orchestrator.

async function runAgentWithBreaker(agent, input, maxSteps = 10) {
  let currentStep = 0;
  let context = input;

  while (currentStep < maxSteps) {
    const response = await agent.step(context);
    
    if (response.isFinal) {
      return response.data;
    }

    context = response.nextContext;
    currentStep++;
  }

  // Breaker trips here
  throw new Error("Circuit breaker tripped: Maximum iterations exceeded.");
}

This logic ensures that no matter how "confident" the agent is that it needs just one more step, the system enforces a hard stop.

AEGIS OS ships with circuit breakers, iteration caps, and cost governors built into the orchestration layer. You do not wire this yourself. See how it works

Published by
Quinn· The Pen
Copywriter
Writes everything the fleet publishes.