Vector and Embedding Weaknesses
- Memory / State
- Tools & External Data
The embedding layer that decides what the model sees — vector stores, similarity search, semantic caches — is exploited through its geometry: poisoned vectors, cross-tenant inference, inversion, or jammed retrieval.
What it is
Any application that turns text, images, code, or audio into vectors and uses similarity search to decide what reaches the prompt makes the embedding layer part of its trust boundary — retrieval-augmented generation most familiarly, but equally vector-backed agent memory, semantic caches, and deduplication. These weaknesses are distinct from prompt injection: they exploit the geometry of the embedding space, and many succeed when the retrieved content carries no instruction at all. OWASP's frame is compact: poisoning makes the system wrong, inversion makes it leak, jamming makes it silent, and access-control failure makes it indiscriminate.
Two facts carry the operational weight. Similarity search often runs across the whole index before an application-layer access filter applies, so result counts, scores, and timing leak what other tenants hold. And modern inversion reconstructs substantial source text from stored vectors, so an "embeddings-only" leak is a source-document leak.
A multi-agent system that shares one vector index or memory store across agents or tenants multiplies both the poisoning surface and the disclosure surface: anything one agent's ingestion pipeline embeds becomes retrievable by every other query against that index unless access is partitioned inside the search itself.
Kinds
- Cross-tenant leakage via shared search
- The access decision runs after the embedding-space search, so probing queries reveal the existence and topic of another tenant's documents.
- Embedding inversion
- An attacker reconstructs source content from exported, backed-up, or exposed vectors.
- Retrieval-time poisoning
- Content crafted so its embedding lands near a target query is retrieved and fed to the model as trusted context.
- Retrieval jamming
- A "blocker" document engineered to be retrieved for a query makes the model refuse or claim it lacks information — an availability attack with no malicious instruction in it.
- Semantic cache poisoning
- Content tuned to sit just across a similarity threshold serves attacker text to every equivalent query, or makes legitimate content get dropped as a duplicate.
Attack scenarios
In a multi-tenant support system, one tenant's planted document ranks highly for another tenant's unrelated query, leaking cross-tenant content into the answer.
Hidden instructions in a resume
An attacker submits a resume with instructions in white-on-white text; a RAG-based screening pipeline ingests it unfiltered and later follows the hidden instruction when queried about the candidate.
Cross-tenant probing
In a shared vector database filtered at the application layer, one tenant's probing queries reveal the existence and topic of another tenant's content through result counts and score gaps.
Inverted vector backup
A misconfiguration exposes a vector-database backup. The incident is rated low-severity because "only the embeddings leaked" — until zero-shot inversion reconstructs the customer conversations behind them.
Mitigations
- Scope inside the query
- Enforce tenant and chunk-level access inside the index search, validated server-side, and use separate indexes per tenant or trust tier for sensitive workloads, per Semantic / Vector / Graph Memory.
- Validate and track provenance before ingestion
- Normalize content (strip zero-width characters, hidden text, homoglyphs), record source and trust tier for every embedding, and vet the embedding model itself — the Integrator's role for retrieval.
- Deny the oracle
- Never return raw similarity scores to clients, rate-limit embedding and search endpoints, and flag new vectors that sit unusually close to many common queries.
- Treat vectors as the documents they encode
- Encrypt embeddings at rest, delete them with their source, protect backups at the source's sensitivity tier, and keep immutable logs of what was retrieved and by whom.