Unbounded Consumption
- Input / Prompt
- Tools & External Data
The application allows excessive or uncontrolled resource usage — inference calls, input, output, and reasoning tokens, tool invocations — enabling denial-of-service, runaway cost, or model extraction.
What it is
Unbounded consumption is any design that lets excessive or uncontrolled LLM inference happen without a limit ever tripping — in request volume, input or output size, or computational cost. It spans denial-of-service, denial-of-wallet through runaway pay-per-use cost, and model extraction, where an attacker collects enough input/output pairs to clone the model's behavior.
Its defining property is cost asymmetry: an attacker triggers expensive computation at negligible cost to themselves. The entry rose four places in 2026 because that asymmetry keeps widening — reasoning models with loosely bounded thinking budgets, multimodal inputs that expand into many tokens, and tool protocols such as MCP that turn one request into a cascade of downstream operations. Request-rate limiting alone no longer suffices.
A multi-agent system multiplies the ways one event becomes unbounded. Agents that retry, re-delegate, or re-query each other can turn one transient error into an unbounded fan-out of inference calls, and a long-lived session that re-processes its growing context on every turn stays under every per-request limit while its total cost climbs.
Kinds
- Flooding and denial of wallet
- Variable-length input floods or a high volume of requests exhaust capacity or run up an unsustainable pay-per-use bill.
- Reasoning-loop exhaustion
- Short, benign-looking prompts push a reasoning model into prolonged or non-terminating thinking, bypassing input-size filters entirely.
- Agent–tool loops and fan-out
- A malicious or badly designed tool drives an agent into recursive calls, or one task spawns hundreds of tool calls.
- Model extraction
- Systematic querying, accelerated by exposed log-probabilities, collects enough output to train a functional equivalent of the target model.
Attack scenarios
A group of agents autonomously re-queries each other in a retry loop after a transient tool failure, multiplying one failed call into a cost spike before any budget check trips.
Request flood
An attacker sends a high volume of requests to the LLM API, exhausting computational resources until the service is unavailable to legitimate users.
Recursive tool from a repository
An attacker publishes a tool — for instance as a skill in an open-source repository — that instructs any agent using it to perform recursive tasks, so every developer who adopts it inherits runaway token consumption.
Agent retry storm
A group of agents autonomously re-queries each other after a transient tool failure, multiplying one failed call into a cost spike before a budget check catches it.
Functional model replication
An attacker uses the target model's own API output as synthetic training data and fine-tunes a separate model into a functional equivalent.
Mitigations
- Cap runs with a hard kill switch
- Token / Cost Tracking enforces a non-overridable ceiling on spend and call count per key, user, and run that halts inference — not an alert a fast workload outpaces.
- Put circuit breakers on every agent
- Enforce step limits, recursion-depth limits, time limits, and per-run cost ceilings, and detect loops by hashing state — the discipline named in the Hallucinated Routing and Unbounded Loops anti-pattern.
- Limit by tokens, not only requests
- Rate-limit tokens per minute and estimated cost per request, and reject oversized inputs with a pre-flight token estimate before inference begins.
- Monitor tool interactions
- Baseline normal tool behavior and flag sessions whose consumption diverges from it without a clear end state.