Security — the attack surface of a multi-agent system.
How multi-agent systems get attacked, and which defenses close each gap. Two OWASP catalogs — the LLM Top 10 and the Agentic Threats — mapped to the guardrails, least-privilege, and human gates this project already documents, plus the draft Agentic Skills Top 10 for the skill layer.
- 5
- Attack surfaces
- 27
- Threats
- 2
- OWASP catalogs
- 10
- ASI patterns
Six ways a MAS amplifies risk a single call never carries.
Adding agents does not add attack surface linearly. Several risks below exist only once agents coordinate, and each raises the cost of one compromise well past anything a lone LLM call could expose.
- 01
Blast radius
A single LLM call fails in isolation. A compromised agent can pass its corrupted state to every agent that trusts its output without re-verifying it, so one exploited hand-off cascades into a system-wide failure.
- 02
Agent collusion
Two or more compromised or subtly misaligned agents can coordinate — through shared memory, negotiated protocol messages, or repeated interaction — to manipulate a decision in a way no individual agent's output would flag as anomalous.
- 03
Identity sprawl
One agent needs one identity. A dozen agents, each holding its own tool credentials and delegated permissions, turn identity and access management into a combinatorial problem — every added identity is one more credential to spoof or over-provision.
- 04
Coordination failures in dynamic environments
The routing, voting, or consensus logic that lets agents adapt to changing conditions is rarely tested against adversarial conditions, and can break down — or be driven into breaking down — once an adversary controls part of that variance.
- 05
Decision-lineage & auditability gaps
No single agent holds the full reasoning trace behind a multi-agent decision, so reconstructing why the system acted means stitching together partial, differently-shaped traces from every agent that touched it.
- 06
Man-in-the-middle on inter-agent channels
The messages agents exchange — hand-offs, blackboard writes, negotiated votes — travel over a channel a single-LLM application never has, and an attacker who intercepts or alters it changes what one agent believes another said.
All 27 threats, located on the architecture.
Before the index lists them, the threat map shows where each threat enters — from untrusted input through the agent runtime to the triggered action. Hover a threat to see where it lands and which further surfaces it touches.
The threats and defenses on this page map directly to the OWASP GenAI Security Project's published catalogs. Read the source documents:
27 threats, ordered by attack surface.
Rather than two separate catalogs, every entry of the OWASP LLM Top 10 and the Agentic Threats (T1–T17) is filed under the surface it enters. Threats spanning several surfaces appear more than once — each row keeps its origin tag (LLM / Agentic) and its other surfaces.
Input / Prompt
6 threatsEverything the system reads — user prompts and untrusted content an agent ingests — where injection enters.
- LLM01Prompt InjectionInput the model reads — a prompt, retrieved content, a tool result, an image, or persistent memory — alters its behavior in ways the developer did not intend.
- LLM06Unbounded ConsumptionThe application allows excessive or uncontrolled resource usage — inference calls, input, output, and reasoning tokens, tool invocations — enabling denial-of-service, runaway cost, or model extraction.
- LLM08Hidden Context ExposureAn attacker extracts, infers, or reconstructs hidden context — the system prompt, developer instructions, retrieved policy text, or tool schemas — in a way that increases their capability.
- T4Resource OverloadAttackers deliberately exhaust an agent's computational, memory, or external-service capacity — including self-triggered task spawning and multi-agent coordination — to degrade performance or cause failure.
- T6Intent Breaking & Goal ManipulationAttackers exploit the lack of separation between data and instructions to alter an agent's planning, reasoning, or self-evaluation, overriding its intended objective — an extension of prompt injection into long-horizon goal state.
- T14Human Attacks on Multi-Agent SystemsAdversaries exploit inter-agent delegation, trust relationships, and workflow dependencies — rather than attacking a single agent directly — to escalate privilege or manipulate AI-driven operations across the system.
Memory / State
6 threatsPersistent state and retrieved context, where poisoned memory biases every later decision.
- LLM01Prompt InjectionInput the model reads — a prompt, retrieved content, a tool result, an image, or persistent memory — alters its behavior in ways the developer did not intend.
- LLM02Sensitive Information DisclosureConfidential data leaves through a channel nobody authorized — the answer, but also a tool argument, a reasoning trace, a log, an embedding, or a measurable property of inference.
- LLM05Data and Model PoisoningData or model artifacts are durably corrupted — in training, fine-tuning, embedding, retrieval corpora, or distribution — so the system still looks functional but behaves as an attacker wants.
- LLM09Vector and Embedding WeaknessesThe embedding layer that decides what the model sees — vector stores, similarity search, semantic caches — is exploited through its geometry: poisoned vectors, cross-tenant inference, inversion, or jammed retrieval.
- T1Memory PoisoningAttackers exploit an agent's reliance on short- or long-term memory to inject malicious or false data, corrupting future decisions, bypassing security checks, or escalating privilege via memory recall.
- T17Supply Chain CompromiseA compromised component — a model, adapter, library, tool, MCP server, prompt template, or build environment — is drawn into the agent, letting an attacker manipulate its actions, exfiltrate data, or run arbitrary code without ever interacting with the agent directly.
Tools & External Data
11 threatsThe tools and external data an agent calls, where misuse and poisoned results turn reach into risk.
- LLM03Excessive AgencyAn agent is granted more autonomous capability — tools, permissions, or unsupervised action — than its task requires, so a model error or manipulation can act, not merely answer wrong.
- LLM04Supply ChainA compromised component — a tampered model, a malicious adapter, a hijacked conversion step, a poisoned dataset, or a vulnerable package — enters the system or is swapped where an artifact is promoted into a trusted environment.
- LLM05Data and Model PoisoningData or model artifacts are durably corrupted — in training, fine-tuning, embedding, retrieval corpora, or distribution — so the system still looks functional but behaves as an attacker wants.
- LLM06Unbounded ConsumptionThe application allows excessive or uncontrolled resource usage — inference calls, input, output, and reasoning tokens, tool invocations — enabling denial-of-service, runaway cost, or model extraction.
- LLM09Vector and Embedding WeaknessesThe embedding layer that decides what the model sees — vector stores, similarity search, semantic caches — is exploited through its geometry: poisoned vectors, cross-tenant inference, inversion, or jammed retrieval.
- T2Tool MisuseAttackers manipulate an agent into abusing its already-granted tools through deceptive prompts, chaining otherwise-legitimate tool calls into an unauthorized sequence while staying within its nominal permissions.
- T3Privilege CompromiseAttackers exploit mismanaged roles, overly broad permissions, or dynamic and inherited privilege to escalate an agent's access beyond its intended scope.
- T4Resource OverloadAttackers deliberately exhaust an agent's computational, memory, or external-service capacity — including self-triggered task spawning and multi-agent coordination — to degrade performance or cause failure.
- T11Unexpected RCE and Code AttacksAttackers exploit an agent's code-generation or code-execution capability to run unsafe or malicious code, escalate privilege, or compromise the host system directly.
- T16Insecure Inter-Agent Protocol AbuseAttackers exploit weaknesses in the coordination protocols agents speak — chiefly MCP and A2A — to bypass consent checks, hijack a protocol transition, or corrupt shared context, turning the connective tissue of a multi-agent system into an actuation path.
- T17Supply Chain CompromiseA compromised component — a model, adapter, library, tool, MCP server, prompt template, or build environment — is drawn into the agent, letting an attacker manipulate its actions, exfiltrate data, or run arbitrary code without ever interacting with the agent directly.
Inter-Agent Communication
10 threatsMessages between agents, where a poisoned or spoofed peer propagates compromise across the system.
- LLM07MisinformationThe model produces incorrect, incomplete, or misleading content that looks credible enough to drive a human decision, an automated workflow, or an agent action.
- T3Privilege CompromiseAttackers exploit mismanaged roles, overly broad permissions, or dynamic and inherited privilege to escalate an agent's access beyond its intended scope.
- T5Cascading Hallucination AttacksAn agent's tendency to generate plausible-but-false content is exploited so the fabrication propagates and amplifies through memory, self-reflection, or inter-agent communication rather than staying contained to one response.
- T6Intent Breaking & Goal ManipulationAttackers exploit the lack of separation between data and instructions to alter an agent's planning, reasoning, or self-evaluation, overriding its intended objective — an extension of prompt injection into long-horizon goal state.
- T7Misaligned & Deceptive BehaviorsAn agent executes harmful, disallowed, or self-preserving actions while outwardly maintaining the appearance of compliance, exploiting the gap between stated and actual behavior.
- T9Identity Spoofing & ImpersonationAttackers exploit weak or missing authentication to impersonate an agent, user, or service, gaining unauthorized access or action while appearing legitimate.
- T12Agent Communication PoisoningAttackers manipulate inter-agent communication channels to inject false information, misdirect decisions, or corrupt shared knowledge across a multi-agent system, extending static data poisoning to transient, in-flight coordination traffic.
- T13Rogue Agents in Multi-Agent SystemsA malicious or compromised agent operates outside its intended boundaries inside a multi-agent architecture, exploiting inter-agent trust to manipulate decisions, corrupt data, or execute unauthorized actions undetected.
- T14Human Attacks on Multi-Agent SystemsAdversaries exploit inter-agent delegation, trust relationships, and workflow dependencies — rather than attacking a single agent directly — to escalate privilege or manipulate AI-driven operations across the system.
- T16Insecure Inter-Agent Protocol AbuseAttackers exploit weaknesses in the coordination protocols agents speak — chiefly MCP and A2A — to bypass consent checks, hijack a protocol transition, or corrupt shared context, turning the connective tissue of a multi-agent system into an actuation path.
Output / Actuation
11 threatsWhat the system emits or actuates, where unvalidated output and excessive agency cause real-world harm.
- LLM02Sensitive Information DisclosureConfidential data leaves through a channel nobody authorized — the answer, but also a tool argument, a reasoning trace, a log, an embedding, or a measurable property of inference.
- LLM03Excessive AgencyAn agent is granted more autonomous capability — tools, permissions, or unsupervised action — than its task requires, so a model error or manipulation can act, not merely answer wrong.
- LLM07MisinformationThe model produces incorrect, incomplete, or misleading content that looks credible enough to drive a human decision, an automated workflow, or an agent action.
- LLM08Hidden Context ExposureAn attacker extracts, infers, or reconstructs hidden context — the system prompt, developer instructions, retrieved policy text, or tool schemas — in a way that increases their capability.
- LLM10Improper Output HandlingModel-generated content is passed downstream — into a shell, a query, a browser, a terminal, a deployment pipeline, another agent — without validation or encoding, so the model's text becomes an execution path.
- T5Cascading Hallucination AttacksAn agent's tendency to generate plausible-but-false content is exploited so the fabrication propagates and amplifies through memory, self-reflection, or inter-agent communication rather than staying contained to one response.
- T7Misaligned & Deceptive BehaviorsAn agent executes harmful, disallowed, or self-preserving actions while outwardly maintaining the appearance of compliance, exploiting the gap between stated and actual behavior.
- T8Repudiation & UntraceabilityAgents act autonomously without sufficient logging or forensic traceability, so decisions and actions cannot be attributed or reconstructed after the fact.
- T10Overwhelming Human-in-the-LoopAttackers exploit human oversight dependencies by flooding reviewers with excessive intervention requests, inducing decision fatigue and rushed, less-scrutinized approvals.
- T11Unexpected RCE and Code AttacksAttackers exploit an agent's code-generation or code-execution capability to run unsafe or malicious code, escalate privilege, or compromise the host system directly.
- T15Human ManipulationAttackers exploit the trust a human user places in an agent's outputs to influence the human's decisions or actions, without the human realizing they are being misled.
The skill layer
OWASP Agentic Skills Top 10
Between an agent and its tools sits the skill layer: packaged instruction files and scripts an agent loads to run a workflow. MCP defines how the model talks to tools; a skill defines what the agent is made to do with them. Skills sit on the tools and input surfaces above, and a third, still-draft OWASP list covers their risks. Each risk links to the patterns that guard it and the threats it is a case of.
DraftAn OWASP Incubator project; version 1.0 is still in public review. Listed as a draft vocabulary — the defenses and threats are this project's own mapping.
| Risk | Severity | Guarded by | Related threats |
|---|---|---|---|
| AST01Malicious SkillsA skill that looks legitimate hides a payload — in its scripts or in its instruction text — and runs it with the host agent's permissions. | Critical | Tool Registry, Sandbox Execution | T17 Supply Chain Compromise, LLM04 Supply Chain |
| AST02Supply Chain CompromiseRegistries without provenance let attackers mass-upload skills, take over maintainer accounts, poison nested dependencies, or turn repository config files into execution paths. | Critical | Tool Registry, Audit Trail | T17 Supply Chain Compromise, LLM04 Supply Chain |
| AST03Over-Privileged SkillsA skill holds more file, network, shell, or credential access than its function needs, so an injected instruction can use permissions the task never required. | High | Least Privilege Agent, Permission-scoped Tools | T3 Privilege Compromise, LLM03 Excessive Agency |
| AST04Insecure MetadataName, description, declared permissions, and risk tier are attacker-controlled and rarely validated: a skill impersonates a brand, understates what it does, or exploits an unsafe parser at load time. | High | Tool Registry, Output Validation / Schema Enforcement | T9 Identity Spoofing & Impersonation, T17 Supply Chain Compromise |
| AST05Untrusted External InstructionsA skill points the agent at external documentation fetched at runtime; that text becomes part of the skill's instructions and can change after review. | High | Integrator, Tool Registry | LLM01 Prompt Injection, T17 Supply Chain Compromise |
| AST06Weak IsolationSkills run in the host agent's own security context — full file system, shell, and network — because sandboxing is missing or off by default. | High | Sandbox Execution, Least Privilege Agent | T11 Unexpected RCE and Code Attacks |
| AST07Update DriftInstalled skills are neither pinned nor verified on update, so known-vulnerable versions stay deployed or a malicious "patch" arrives silently. | Medium | Tool Registry | T17 Supply Chain Compromise, LLM04 Supply Chain |
| AST08Poor ScanningCode scanners miss payloads written as plain-language instructions, and scanners that use an LLM as judge can be prompt-injected themselves. | Medium | LLM-as-Judge, Multimodal Guardrails | T6 Intent Breaking & Goal Manipulation |
| AST09No GovernanceNo inventory, approval, audit trail, or revocation exists for installed skills, so a compromise is neither seen nor contained. | Medium | Audit Trail, Least Privilege Agent | T8 Repudiation & Untraceability |
| AST10Cross-Platform ReuseA skill ported to another agent platform silently loses security metadata — permissions, risk tier, signature — the target format cannot express. | Medium | Tool Registry | T17 Supply Chain Compromise |
Agentic AI Patterns
The OWASP Agentic Security Initiative's shared vocabulary for threat-modeling conversations. Each maps to a pattern this project already covers in depth.
Reflective Agent
Agents that iteratively evaluate and critique their own outputs to enhance performance.
AI code generators that review and debug their own outputs, like Codex with self-evaluation.
Maps to: Reflexion →Task-Oriented Agent
Agents designed to handle specific tasks with clear objectives.
Automated customer-service agents for appointment scheduling or returns processing.
Maps to: ReAct →Hierarchical Agent
Agents organized in a hierarchy, managing multi-step workflows or distributed control systems.
Project-management systems where higher-level agents oversee task delegation.
Maps to: Hierarchical Supervisor →Coordinating Agent
Agents facilitate collaboration, coordination, and tracking, ensuring efficient execution.
A coordinator assigns subtasks to specialists in an AI-powered DevOps workflow: one plans deployments, another monitors performance, a third handles rollbacks.
Maps to: Orchestrator-Workers →Distributed Agent Ecosystem
Agents interact within a decentralized ecosystem, often in IoT or marketplaces.
Autonomous IoT agents managing smart-home devices, or a marketplace with buyer and seller agents.
Maps to: Swarm / Contract-Net →Human-in-the-Loop Collaboration
Agents operate semi-autonomously with human oversight.
AI-assisted medical-diagnosis tools that recommend but let doctors make the final decision.
Maps to: HITL Gate →Self-Learning and Adaptive Agents
Agents adapt through continuous learning from interactions and feedback.
Co-pilots that adapt to user interactions over time, learning from feedback.
Maps to: Skill-Build / Reflector →RAG-Based Agent
Agents use Retrieval-Augmented Generation to draw on external knowledge sources dynamically.
Agents performing real-time web browsing for research assistance.
Maps to: Agentic RAG →Planning Agent
Agents autonomously devise and execute multi-step plans to achieve complex objectives.
Task-management systems organizing and prioritizing tasks by user goals.
Maps to: Plan-and-Execute →Context-Aware Agent
Agents dynamically adjust their behavior and decision-making based on the context in which they operate.
Smart-home systems adjusting settings based on user preferences.
Maps to: Virtual Context Management →
Security