Prompt Injection

Attack surface
  • Input / Prompt
  • Memory / State

Input the model reads — a prompt, retrieved content, a tool result, an image, or persistent memory — alters its behavior in ways the developer did not intend.

What it is

A model makes no architectural distinction between instructions and data — both are tokens on the same stream — so there is no clean equivalent to a parameterized query, and no reliable prevention mechanism exists today. Prompt injection is any unintended change in the model's behavior driven by what it reads, whether or not that input is human-readable, comes from the user, or is visible on screen. Jailbreaking is the subset aimed at the model's safety protocols.

Three deployment properties widen the exposure. The system prompt, user input, retrieved documents, tool output, and memory are pooled into one context with no enforced trust boundary. An injection that writes to long-term memory or a retrieval corpus taints every later session that reads it. And once the model's output drives tool calls, the blast radius reaches whatever those tools can reach, while tool output flows back into the context. A multi-agent system adds every peer message to that list of channels.

Defense is therefore architectural rather than interceptive. Most high-impact incidents became severe because the injection landed in a system whose tools, scopes, or rendering let the compromised model act at the user's privilege level — which is why this entry pairs with Excessive Agency (LLM03).

Kinds

Direct injection
The user, or an attacker with the user's access path, supplies input that overrides the system's instructions — intentionally, as a jailbreak, or by accident, when pasted content happens to carry conflicting instructions.
Indirect
The model ingests instructions embedded in external content — a web page, an email, a retrieved passage, a tool or MCP-server response, an issue title. The source may be untrusted, semi-trusted (a public bug tracker), or even trusted (the developer's own repository, reached through a low-privilege upstream channel).
Multimodal and encoded
Instructions ride in a non-text channel, such as sub-perceptual image or audio perturbations, or in an encoding a filter never saw: invisible Unicode characters, Base64, or a low-resource language.

Attack scenarios

In a multi-agent system

A web-search tool's result contains a hidden instruction that a research agent follows as if it came from the user, silently changing the agent's next action.

Support bot bypass

An attacker submits crafted input to a customer-support chatbot that overrides its guidelines, coaxes it into querying private data, and drives it to send emails on the attacker's behalf.

Poisoned webpage summary

An LLM asked to summarize a webpage processes a hidden instruction embedded in the page, which causes it to insert a Markdown image whose URL silently exfiltrates the conversation.

Split payload in a resume

An attacker splits a malicious instruction across several resume fields so no single field looks malicious; the model recombines the fragments when it reads the whole document and skews its evaluation.

Poisoned issue, privileged agent

An attacker plants text in a public GitHub issue. A developer's MCP-connected agent reads it under the developer's own credentials and exfiltrates private repositories — the attacker never touches the backend; the agent performs the privileged action.

Mitigations

Hold capability outside the model
Keep credentials and state-changing capability in application code, grant least privilege per operation, and route privileged calls through a deterministic policy check, per Least Privilege Agent and Permission-scoped Tools.
Budget agent capabilities
Apply the Rule of Two as a floor: an agent that at once processes untrusted input, reaches sensitive data, and changes state or communicates externally needs per-action approval through a HITL Approval Gate.
Separate and filter every channel
Pass external content through a provenance-labeled channel distinct from the user's instructions, strip invisible Unicode at every ingest and render boundary, and filter every modality, not text alone, per Multimodal Guardrails.
Treat memory writes as privileged
Log the prompt that caused a memory write, and require approval before instruction-bearing memories persist across sessions.
Test against adaptive attackers
Red-team with the deployed defense disclosed to the testers: attack-success rates near zero against static attacks routinely rise above 90% once the attacker adapts.

Security

Where to next

Search

Search patterns, frameworks, and pages.