Sensitive Information Disclosure
- Output / Actuation
- Memory / State
Confidential data leaves through a channel nobody authorized — the answer, but also a tool argument, a reasoning trace, a log, an embedding, or a measurable property of inference.
What it is
Sensitive information disclosure covers any confidential content an LLM-integrated system reveals through a channel it was not meant to cross — personal and health data, credentials, trade secrets, model weights. The channel is not only the answer: tool-call arguments, reasoning traces, retrieved chunks, logs, telemetry, embeddings, and measurable properties such as token length or latency are all disclosure surfaces and deserve the same classification and redaction rules.
Two structural failures drive most incidents. The first is oversharing upstream: unscoped drives and legacy permissions feed a retrieval index with data the model then retrieves exactly as designed, so the fix lies in the data surface, not the model. The second is persistence: once data has influenced weights, embeddings, or adapters, it stays extractable after the source is deleted.
A multi-agent system widens the exposure with every surface it adds — memory pooled across sessions or tenants, a retrieval index fed by several agents, hand-offs that pass raw context onward, and a shared observability pipeline that logs every agent's prompts and traces by default.
Kinds
- Training-time
- A model, fine-tune, or LoRA adapter memorizes corpus content and later reproduces it; narrow adapters memorize rare examples with high fidelity.
- Inference-time
- The model discloses live context — the system prompt, retrieved chunks, files, tool output, memory, another session's data — often because a summary or extraction surfaces more than was asked.
- Pipeline-time
- Fine-tuning, distillation, synthetic-data generation, and observability tooling move sensitive data into derived artifacts.
- Observation-time
- An adversary infers facts from token length, latency, log-probabilities, or cache-hit behavior without ever receiving content.
Attack scenarios
A support agent with memory shared across tenants surfaces one customer's account details while answering a different customer's question.
Memorized data at scale
Divergence prompts make a production model emit memorized personal data and live credentials at scale.
Traces in the monitoring tool
Extended-thinking traces logged verbatim to a shared monitoring project expose retrieved personal data to hundreds of engineers, while the visible answer stays sanitized.
Cross-client retrieval
A retrieval index shared across a legal firm's clients synthesizes one client's privileged strategy into another client's answer.
"Embeddings-only" backup
A leaked vector backup is first rated low-risk, then reclassified as a source-document breach once the vectors are inverted back to text.
Mitigations
- Minimize what enters context
- Classify and scrub corpora at ingest, and send only the fields a task requires to an external provider.
- Authorize before retrieval
- Enforce document- and chunk-level authorization inside the index query, never as a filter after retrieval, per Semantic / Vector / Graph Memory and Least Privilege Agent.
- Treat every channel as output
- Classify and redact reasoning traces and tool arguments like the final answer, check responses with Output Validation / Schema Enforcement, and use trained classifiers — regex alone fails on encoded and cross-lingual output.
- Log what crossed which boundary
- Record disclosure paths in the Audit Trail, and scrub traces before they reach an unrestricted observability platform.