How to protect your data from frontier AI labs

Short answer

Whether a frontier lab can train on your data is determined by the tier and contract you are on, not by the model. Consumer tiers have historically had far weaker terms than enterprise agreements, which commonly exclude training on customer content and offer retention controls. The durable protections are contractual — no-training terms, defined retention, residency, sub-processor disclosure — combined with architectural ones: minimize what you send, redact what is not needed, and keep the crown-jewel data out entirely.

4 min readUpdated 2026-09-28Problems We Solve

The concern is legitimate and frequently framed imprecisely. "Will the AI steal our data" is not quite the right question. The right questions are: what is retained, for how long, who can see it, is it used to improve models, and what happens if that changes.

What actually happens to a prompt

When you call a model API, your prompt goes to the provider's infrastructure. From there, four things vary by contract:

Training use. Whether your content is used to improve their models. Enterprise and business agreements commonly exclude this; consumer tiers historically have not, and defaults differ across providers.

Retention. How long the content is stored. Often 30 days for abuse monitoring by default, with zero-retention options available on request for qualifying use cases.

Human review. Whether staff may read flagged content for safety purposes.

Sub-processors. Which other parties touch the data — commonly the underlying cloud provider.

These change over time and differ by provider and product. Read the current data-processing addendum for the specific tier you are on rather than relying on a general impression — including any summary here.

The most common serious mistake is assuming enterprise terms apply because your company has an enterprise account, while individual employees use personal consumer accounts for the same work. See employees pasting company data into chatbots.

Contractual controls that matter

Get these explicit, in writing, for each provider:

  1. No training on customer content, stated unambiguously and covering both inputs and outputs.
  2. Defined retention, with zero-retention where available and justified.
  3. Data residency if you have jurisdictional requirements.
  4. Sub-processor list and change notification.
  5. Confidentiality and security commitments, plus a current independent audit report.
  6. Deletion on termination, with a defined timeframe.
  7. Notice of material changes to data handling, with an exit right.

That last one matters more than it looks. Terms change. An exit right on material change is what keeps a good agreement good.

Architectural controls, which you control entirely

Contracts allocate liability. Architecture prevents exposure. Use both.

Send less. The strongest protection is not sending the data at all. Most tasks need a fraction of the record. Ship the paragraph, not the document; the field, not the row.

Redact identifiers you do not need. Names, account numbers, and contact details are rarely required for summarization, classification, or extraction of other fields. Tokenize them out and restore afterwards. Note that automated redaction is imperfect — treat it as risk reduction, not as making data non-sensitive.

Classify before you route. Define tiers of sensitivity and which model deployment each may use. Public and internal content can go to a commercial API under good terms; restricted content may require a private deployment; crown-jewel data may not leave your boundary at all.

Use private deployments for sensitive workloads. Major clouds offer model hosting inside your own tenancy where content does not leave your environment. Cost and capability tradeoffs are real but often worth it.

Self-host open-weight models for the most sensitive work. Maximum control, real operational burden — see build vs buy.

Never put secrets in prompts. A system prompt is not a security boundary. Credentials, keys, and internal URLs in prompts will eventually surface in output or logs.

Secure the derived artifacts. Your vector store, caches, and traces hold copies of the same content, usually with weaker controls. Embeddings are recoverable — treat them as sensitive. See how AI systems leak data.

The competitive-advantage question

Beyond compliance there is a strategic concern: does using a frontier model on your proprietary data erode your advantage?

Under a no-training agreement, your content is not improving a competitor's model. But consider two further points. First, your prompts encode how you work — the questions you ask reveal your method. Second, if your entire AI capability is a thin wrapper over a commodity API, your competitors can replicate it, regardless of data.

The defensible position is that your advantage lives in your proprietary data, your evaluation set, your retrieval strategy, and your integration — none of which transfer to the model provider, and none of which a competitor can copy by calling the same API.

Frequently asked questions

Do AI companies train on business API data?

Major providers' enterprise and standard API agreements generally exclude training on customer content, while consumer chat products have historically had different defaults. This is contractual and varies by provider and tier, so verify against the current data-processing terms for your specific plan rather than relying on general statements.

Is zero data retention available?

Several providers offer zero-retention configurations for qualifying API use, often requiring an application or specific plan. It typically disables features that depend on stored state. Ask your provider directly — it is frequently available and not advertised prominently.

Is it safer to use a cloud provider's hosted model?

Often, yes, for governance reasons. Hosting a model inside your existing cloud tenancy keeps content within an environment you have already assessed, under your existing agreements and controls. Capability may lag the direct provider slightly, and cost differs, but the compliance story is much simpler.

What about data we already sent?

Request deletion under your agreement, and check retention terms for what remains. For content sent through consumer accounts before a policy existed, your options are limited — which is the reason to establish the sanctioned path and the policy now rather than after the next incident.

Does self-hosting eliminate the risk?

It eliminates third-party disclosure and replaces it with your own operational risk — patching, access control, and securing the model endpoint. It is the right answer where data sensitivity requires it, and overkill for most workloads under good contractual terms.

Guardian Robotics is an AI consultancy.

We build the pipelines, agents, and automation this article describes — for commercial teams and federal agencies alike.