Intent Breaking & Goal Manipulation

Attack surface
  • Input / Prompt
  • Inter-Agent Communication

Attackers exploit the lack of separation between data and instructions to alter an agent's planning, reasoning, or self-evaluation, overriding its intended objective — an extension of prompt injection into long-horizon goal state.

What it is

Intent breaking and goal manipulation exploits the lack of separation between data and instructions in an agent's planning loop: prompt injections, compromised data sources, or malicious tool output alter its planning, reasoning, or self-evaluation, letting an attacker override the intended objective, redirect decision-making, or force unauthorized actions. Adaptive, reasoning-heavy architectures — a ReAct-style planner that re-evaluates its own goal at every step — are the most exposed, because each step reasons from whatever state the previous step left behind.

The threat is an extension of Prompt Injection (LLM01) into long-horizon goal state: where a single injected prompt changes one response, a goal-manipulation attack changes what the agent is trying to accomplish across an entire multi-step run, so a single successful injection compounds across every later planning step rather than staying scoped to one output. A related failure mode is agent hijacking (see Tool Misuse), where the redirection happens through data the agent ingests rather than a direct instruction.

Kinds

Direct injection
An attacker instructs the agent outright — bypassing its guardrails — to ignore its original instructions and chain tool executions into an unauthorized sequence.
Indirect injection
A maliciously crafted tool result or document smuggles hidden goal-altering instructions that the agent misinterprets as part of its own operational goal.
Gradual goal drift
An attacker incrementally injects subtly modified sub-goals across many turns, so the agent's objective drifts away from the original task while every individual step still looks reasonable.

Attack scenarios

In a multi-agent system

An attacker incrementally injects subtly modified sub-goals into a planning agent's context across several turns, drifting its objective away from the original task while each individual step still looks reasonable.

Chatbot instruction override

An attacker instructs a chatbot to ignore its original instructions and instead chain tool executions to exfiltrate data or send unauthorized emails.

Poisoned tool-output redirection

A maliciously crafted tool output introduces hidden instructions that the agent misinterprets as part of its operational goal, leading to sensitive-data exfiltration.

Reflection-loop paralysis

An attacker triggers infinite or excessively deep self-analysis cycles in the agent, consuming resources and preventing it from making real-time decisions.

Meta-learning corruption

By manipulating an agent's self-improvement mechanisms, an attacker introduces learning patterns that progressively alter its decision-making integrity, enabling unauthorized actions over time.

Documented incidents

Public cases, each checked against its primary source. A research demonstration shows the attack works against real systems; it is not a reported breach.

Disclosed vulnerability · August 2025

Perplexity Comet browser agent hijacked

A user asked the Comet browser agent to summarize a Reddit page. A hidden comment on the page replaced that goal: the agent collected the user's email address and a login code from the open Gmail tab and sent both to the attacker. The agent did not separate the user's instruction from untrusted page content.

Brave, “Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet”
Research demonstration · October 2024

ZombAIs: a computer-use agent joins a botnet

A web page asked Anthropic's computer-use agent to download and run a “support tool”. The agent abandoned its task, downloaded the binary, made it executable, and ran it, connecting the machine to the attacker's command-and-control server.

Johann Rehberger, “ZombAIs: From Prompt Injection to C2 with Claude Computer Use”

Mitigations

Shrink the attack surface
Restrict tool access to the minimum a task needs, and validate every AI output before it's treated as a plan or reused downstream.
Validate goal consistency
Check every planning step with Output Validation / Schema Enforcement, and track how often an agent requests to change its own goals — a repeated pattern is itself a signal of manipulation.
Bound self-reinforcement
A Controller applies goal-alignment monitoring and behavioral constraints so an agent cannot self-adjust its objectives beyond predefined operational parameters, closing the reflection-loop and meta-learning surface.
Make every plan change traceable and reviewable
The Audit Trail keeps an immutable, cryptographically verifiable log of every planning decision, and a HITL Approval Gate requires human sign-off before a plan change takes effect.
MAESTRO layer(s)

Security

Where to next

Search

Search patterns, frameworks, and pages.