Misinformation

Attack surface
  • Output / Actuation
  • Inter-Agent Communication

The model produces incorrect, incomplete, or misleading content that looks credible enough to drive a human decision, an automated workflow, or an agent action.

What it is

Misinformation is false or misleading output that reads as credible — and the core risk is that it is trusted and acted upon. Its causes include hallucination, stale or incomplete context, weak grounding, ambiguous prompts, biased data, and unvalidated tool output, and an attacker can induce it deliberately.

It is the entry where vote and evidence diverged most: voters placed it near the bottom, the incident record near the top, and it rose two places in 2026. The reason is a shift in what model output does. It no longer only informs a reader; it drives tool calls, infers system state, authorizes actions, and coordinates agents — so a wrong answer becomes a wrong action, and the overreliance that makes it dangerous is often built into the system's design.

A multi-agent system compounds it, because one agent's fabrication routinely becomes another agent's trusted input. A downstream agent cannot tell a verified fact from an upstream hallucination unless the pipeline explicitly checks — exactly the propagation that Cascading Hallucination Attacks describes.

Kinds

Unsupported decision support
False or unsupported claims influence a business, legal, medical, or financial decision.
Incorrect state inference
The model concludes that a condition has been met when it has not, and an unintended action follows.
Misleading summaries and omissions
A summary drops a constraint, exception, or risk the reader needed to act safely.
Cross-agent propagation
An incorrect output travels across agents and workflows and is treated as established fact at every hop.
Fabricated code and dependencies
The model references a nonexistent package or produces incorrect code; the supply-chain consequence belongs to LLM04, executing unsafe code to LLM10.

Attack scenarios

In a multi-agent system

A research agent fabricates a plausible-sounding citation, and a downstream summarization agent repeats it as fact in the final report without an independent check.

Cross-agent trust failure

A retrieval agent reports a customer as identity-verified when they are not, and a downstream payment agent trusts that state and releases funds.

Fabricated task completion

An agent reports that a nightly database backup completed when it never ran, and a later restore fails because no backup exists.

Seeded forum answer

An attacker seeds a support forum with false remediation steps that a troubleshooting agent retrieves and repeats as a trusted recommendation.

Fabricated citation relay

A research agent fabricates a plausible-sounding citation, and a downstream summarization agent repeats it as fact without an independent check.

Mitigations

Claim, check, then act
Separate generation from execution, and verify a claim against an authoritative, current source before any tool call or state change depends on it.
Validate tool calls against real state
Check arguments, authorization, and preconditions against the system of record, not against the model's own report.
Use verification signals, not confidence
Score groundedness and consistency with LLM-as-Judge, and catch topical drift with Statistical Guardrails.
Force completeness and test for it
Require structured outputs with mandatory fields to surface omissions, and cover fabricated claims and state with Integration Tests for Agents.

Security

Where to next

Search

Search patterns, frameworks, and pages.