Most AI data leaks are not exotic. They are ordinary access-control failures in a new architecture where the usual controls were never wired up.
1. Retrieval that ignores permissions
The most common serious flaw in production AI systems.
You index your document corpus for RAG. A user asks a question. The system searches everything, retrieves the most relevant passages, and answers — using documents that user was never entitled to read.
There was no breach in the conventional sense. The permissions existed in the source system; the index simply did not carry them.
The fix is to enforce access control at the retrieval step, filtering candidates by the asking user's entitlements before ranking. That means indexing permission metadata alongside every chunk and keeping it synchronized as permissions change — including revocations, which are the part teams forget.
Test this explicitly. Create a low-privilege test account, ask it questions whose answers live in restricted documents, and see what comes back. A great many deployed systems fail this test, and almost none have been checked.
2. Prompts sent to third parties
Everything in a prompt goes to the model provider. That includes retrieved documents, conversation history, and any customer data you enriched with.
Controls: enterprise agreements with explicit no-training and retention terms, data residency where required, redaction of identifiers that are not needed for the task, and a clear classification of what content may leave your boundary at all.
3. Logs, traces, and observability
You instrumented the pipeline for debugging. Those traces contain full prompts and outputs — which means your observability stack now holds a copy of every sensitive document that passed through, often with looser access controls and longer retention than the source system.
Controls: redact before logging, apply the source system's access controls to traces, set short retention, and treat the trace store as a system of record for compliance purposes.
4. Model output
A model can reveal context it was given, including material retrieved for a different purpose in the same session. Fine-tuned models can regurgitate training examples. A model told a secret in a system prompt will eventually disclose it — system prompts are not a security boundary.
Controls: never put credentials or secrets in prompts, scope each session's context to one purpose, and validate output before display.
5. Agent tools
An agent with an outbound capability — email, HTTP request, file write, message post — is an exfiltration path. Combined with prompt injection, attacker-controlled content can direct the agent to transmit data outward.
Controls: least privilege on tools, allowlisted destinations, human approval on outbound actions, and egress monitoring on what agents actually send.
6. Embeddings
Vector embeddings are frequently treated as opaque. They are not. Research has repeatedly demonstrated that meaningful information — and in some cases substantial verbatim text — can be reconstructed from embeddings.
Controls: secure the vector store to the same standard as the source documents. An embedding is derived sensitive data, not an anonymized artifact.
A practical review checklist
- Does retrieval filter by the asking user's permissions, and does revocation propagate?
- What contractual terms govern data sent to each model provider?
- What do our traces contain, who can read them, and for how long are they kept?
- Can any tool transmit data outside our boundary, and is the destination constrained?
- Is the vector store secured like the source corpus?
- Have we tested with a deliberately low-privilege account?
The checklist version of this review is published as controls GCAS-13 to GCAS-16 in our Controlled Autonomy Standard, each with the evidence we expect.
Frequently asked questions
Does using an enterprise AI plan prevent data leakage?
It addresses one channel — the provider's use and retention of your data — and none of the others. Enterprise terms do nothing about a retrieval layer that ignores permissions, traces that log everything, or an agent that can email data out. Those are your architecture's responsibility.
How do we apply access control to a vector database?
Store permission metadata with each chunk — owning group, sensitivity label, source document ID — and filter on it at query time before ranking. Then keep it in sync with the source system, which is the hard part. Some teams instead partition into separate indexes per security boundary, which is simpler to reason about and harder to get wrong.
Is it safe to fine-tune on sensitive data?
Fine-tuned models can reproduce training examples, so a model trained on sensitive data must be treated as sensitive itself — including access control over who can query it. For most cases where teams consider this, retrieval with proper access control is both safer and more appropriate, since it keeps the data in a system you can apply permissions to.
What about de-identification before sending to a model?
Useful and imperfect. Automated redaction misses identifiers, and de-identified records can often be re-identified when combined with other data. It reduces risk and should not be treated as making data non-sensitive.
Who owns AI data governance?
In practice it works when security, legal, and the engineering team building the pipeline share it, with one named accountable owner. It fails when it is assigned to a committee with no engineering authority, because most of the controls are architectural decisions made in code.
