Unbounded Consumption

Attack surface
  • Input / Prompt
  • Tools & External Data

The application allows excessive or uncontrolled resource usage — inference calls, input, output, and reasoning tokens, tool invocations — enabling denial-of-service, runaway cost, or model extraction.

What it is

Unbounded consumption is any design that lets excessive or uncontrolled LLM inference happen without a limit ever tripping — in request volume, input or output size, or computational cost. It spans denial-of-service, denial-of-wallet through runaway pay-per-use cost, and model extraction, where an attacker collects enough input/output pairs to clone the model's behavior.

Its defining property is cost asymmetry: an attacker triggers expensive computation at negligible cost to themselves. The entry rose four places in 2026 because that asymmetry keeps widening — reasoning models with loosely bounded thinking budgets, multimodal inputs that expand into many tokens, and tool protocols such as MCP that turn one request into a cascade of downstream operations. Request-rate limiting alone no longer suffices.

A multi-agent system multiplies the ways one event becomes unbounded. Agents that retry, re-delegate, or re-query each other can turn one transient error into an unbounded fan-out of inference calls, and a long-lived session that re-processes its growing context on every turn stays under every per-request limit while its total cost climbs.

Kinds

Flooding and denial of wallet
Variable-length input floods or a high volume of requests exhaust capacity or run up an unsustainable pay-per-use bill.
Reasoning-loop exhaustion
Short, benign-looking prompts push a reasoning model into prolonged or non-terminating thinking, bypassing input-size filters entirely.
Agent–tool loops and fan-out
A malicious or badly designed tool drives an agent into recursive calls, or one task spawns hundreds of tool calls.
Model extraction
Systematic querying, accelerated by exposed log-probabilities, collects enough output to train a functional equivalent of the target model.

Attack scenarios

In a multi-agent system

A group of agents autonomously re-queries each other in a retry loop after a transient tool failure, multiplying one failed call into a cost spike before any budget check trips.

Request flood

An attacker sends a high volume of requests to the LLM API, exhausting computational resources until the service is unavailable to legitimate users.

Recursive tool from a repository

An attacker publishes a tool — for instance as a skill in an open-source repository — that instructs any agent using it to perform recursive tasks, so every developer who adopts it inherits runaway token consumption.

Agent retry storm

A group of agents autonomously re-queries each other after a transient tool failure, multiplying one failed call into a cost spike before a budget check catches it.

Functional model replication

An attacker uses the target model's own API output as synthetic training data and fine-tunes a separate model into a functional equivalent.

Mitigations

Cap runs with a hard kill switch
Token / Cost Tracking enforces a non-overridable ceiling on spend and call count per key, user, and run that halts inference — not an alert a fast workload outpaces.
Put circuit breakers on every agent
Enforce step limits, recursion-depth limits, time limits, and per-run cost ceilings, and detect loops by hashing state — the discipline named in the Hallucinated Routing and Unbounded Loops anti-pattern.
Limit by tokens, not only requests
Rate-limit tokens per minute and estimated cost per request, and reject oversized inputs with a pre-flight token estimate before inference begins.
Monitor tool interactions
Baseline normal tool behavior and flag sessions whose consumption diverges from it without a clear end state.

Security

Where to next

Search

Search patterns, frameworks, and pages.