Sensitive Information Disclosure

Attack surface
  • Output / Actuation
  • Memory / State

Confidential data leaves through a channel nobody authorized — the answer, but also a tool argument, a reasoning trace, a log, an embedding, or a measurable property of inference.

What it is

Sensitive information disclosure covers any confidential content an LLM-integrated system reveals through a channel it was not meant to cross — personal and health data, credentials, trade secrets, model weights. The channel is not only the answer: tool-call arguments, reasoning traces, retrieved chunks, logs, telemetry, embeddings, and measurable properties such as token length or latency are all disclosure surfaces and deserve the same classification and redaction rules.

Two structural failures drive most incidents. The first is oversharing upstream: unscoped drives and legacy permissions feed a retrieval index with data the model then retrieves exactly as designed, so the fix lies in the data surface, not the model. The second is persistence: once data has influenced weights, embeddings, or adapters, it stays extractable after the source is deleted.

A multi-agent system widens the exposure with every surface it adds — memory pooled across sessions or tenants, a retrieval index fed by several agents, hand-offs that pass raw context onward, and a shared observability pipeline that logs every agent's prompts and traces by default.

Kinds

Training-time
A model, fine-tune, or LoRA adapter memorizes corpus content and later reproduces it; narrow adapters memorize rare examples with high fidelity.
Inference-time
The model discloses live context — the system prompt, retrieved chunks, files, tool output, memory, another session's data — often because a summary or extraction surfaces more than was asked.
Pipeline-time
Fine-tuning, distillation, synthetic-data generation, and observability tooling move sensitive data into derived artifacts.
Observation-time
An adversary infers facts from token length, latency, log-probabilities, or cache-hit behavior without ever receiving content.

Attack scenarios

In a multi-agent system

A support agent with memory shared across tenants surfaces one customer's account details while answering a different customer's question.

Memorized data at scale

Divergence prompts make a production model emit memorized personal data and live credentials at scale.

Traces in the monitoring tool

Extended-thinking traces logged verbatim to a shared monitoring project expose retrieved personal data to hundreds of engineers, while the visible answer stays sanitized.

Cross-client retrieval

A retrieval index shared across a legal firm's clients synthesizes one client's privileged strategy into another client's answer.

"Embeddings-only" backup

A leaked vector backup is first rated low-risk, then reclassified as a source-document breach once the vectors are inverted back to text.

Mitigations

Minimize what enters context
Classify and scrub corpora at ingest, and send only the fields a task requires to an external provider.
Authorize before retrieval
Enforce document- and chunk-level authorization inside the index query, never as a filter after retrieval, per Semantic / Vector / Graph Memory and Least Privilege Agent.
Treat every channel as output
Classify and redact reasoning traces and tool arguments like the final answer, check responses with Output Validation / Schema Enforcement, and use trained classifiers — regex alone fails on encoded and cross-lingual output.
Log what crossed which boundary
Record disclosure paths in the Audit Trail, and scrub traces before they reach an unrestricted observability platform.

Security

Where to next

Search

Search patterns, frameworks, and pages.