Data and Model Poisoning
- Memory / State
- Tools & External Data
Data or model artifacts are durably corrupted — in training, fine-tuning, embedding, retrieval corpora, or distribution — so the system still looks functional but behaves as an attacker wants.
What it is
Data and model poisoning corrupts a system's learning process rather than a single runtime path: pre-training, fine-tuning, or embedding data — or a model artifact itself — is manipulated to introduce a bias, a vulnerability, or a backdoor, deliberately or through poor data hygiene. Because the model learns the wrong patterns, the fix is rarely a code patch; it can mean revalidating data, retraining, or replacing the model.
Two findings shape the 2026 entry. Scale does not protect: as few as 250 poisoned documents compromised models from 600M to 13B parameters, regardless of dataset size. And safety training does not remove what poisoning planted: a backdoor stays dormant until its trigger fires, a sleeper agent ordinary evaluation never meets. The entry now also absorbs fine-tuning subversion — targeted data that erodes refusal behavior while general accuracy stays intact.
The boundaries are explicit: this entry covers durable corruption. Instructions delivered through retrieved content at inference time belong to Prompt Injection (LLM01), attacks on embedding geometry to Vector and Embedding Weaknesses (LLM09). A multi-agent system raises the stakes, because a shared retrieval index, a shared memory layer, or a feedback loop that retrains on agent interactions is read by every agent that relies on it.
Kinds
- Training and fine-tuning poisoning
- Biased or malicious examples enter pre-training or fine-tuning data — including targeted data that erodes refusals without degrading accuracy.
- Retrieval and memory poisoning
- A falsified document planted in a shared knowledge base, or instructions persisted into agent memory over several sessions, is treated as ground truth by every later reader.
- Backdoor insertion
- A trigger phrase is trained into the model so its behavior stays normal until the trigger appears — hard to detect, because ordinary evaluation never encounters the trigger.
- Poisoned inference artifacts
- A chat template, tokenizer configuration, adapter, or quantization artifact redistributed through a public hub carries trigger-activated instructions or code that runs on load.
Attack scenarios
An attacker seeds a knowledge base shared by several agents with a subtly falsified policy document, so every agent that retrieves from it inherits the same wrong answer.
Shared knowledge-base poisoning
A malicious actor seeds a knowledge base shared by several agents with subtly falsified documents; every agent that retrieves from it inherits and repeats the same wrong answer.
Retraining-loop drift
An attacker submits crafted inputs through an ordinary user interface into an automated retraining loop, slowly drifting the model toward biased or unsafe recommendations.
Trojaned chat template
An attacker modifies a model's chat template with trigger-activated instructions; redistributed through a public hub, the model behaves normally until the trigger appears.
Persistent memory takeover
An attacker injects instructions into an agent's persistent memory across several sessions until the agent prioritizes attacker-controlled logic.
Mitigations
- Track lineage and verify integrity
- Record dataset and model lineage in a bill of materials, sign artifacts, and treat chat templates, tokenizer configurations, and adapters as security-relevant code.
- Partition and permission-scope memory
- Partition retrieval stores per Semantic / Vector / Graph Memory so a poisoned entry in one partition cannot answer queries from an unrelated agent or tenant.
- Validate shared stores and feedback loops
- Screen everything entering a shared index with the Integrator, and put human oversight and rate limits on automated retraining.
- Monitor for drift, probe for triggers
- Watch training loss and live output for the anomalies Statistical Guardrails are built to catch, and red-team for backdoor triggers after every alignment cycle instead of assuming safety training removed them.