LLM API Privacy Due Diligence: 3+ of 13 Train or Improve on Your Data
At least three of the 13 LLM API vendors we reviewed say in their terms that they use customer data to train or improve their services by default: DeepSeek, Kimi and MiniMax. Four more have terms too unclear to call. Those three, plus Baidu Qianfan’s international offering, have no public data processing addendum (DPA), and we found no processor terms in their public terms either. OpenAI, Anthropic, Google’s paid Gemini API tier, xAI, Alibaba Cloud Model Studio (international) and BytePlus do not train on API data by default. The six US, EU and Canadian vendors keep prompts and outputs for 30 to 55 days for abuse monitoring, and only Z.ai states in writing that it stores nothing by default. If you serve US or EU users, plugging in an LLM API usually means adding a processor, and both GDPR and the CCPA require a written contract that limits what that processor can do. If the data involves Americans, the DOJ’s data security rule may apply as well.
Key takeaways
- Six vendors do not train on API data by default: OpenAI, Anthropic, Google (paid Gemini API), xAI, Alibaba Cloud Model Studio (international) and BytePlus. Three use it to train or improve their services by default: DeepSeek, Kimi (international) and MiniMax (international). Only DeepSeek says “train” explicitly; the other two say “develop and improve the Services.” The terms are unclear for Mistral pay-as-you-go, Cohere, Z.ai and Baidu Qianfan (international).
- Only Z.ai’s DPA says API content is not stored by default. OpenAI, Anthropic, Mistral, xAI and Cohere keep it for 30 days by default and Google’s Gemini Developer API for 55 days. Most Chinese vendors’ overseas offerings say only “as long as necessary” or “while the account is active.”
- Where the contract is signed and where inference runs are separate questions. DeepSeek processes and stores data in mainland China. Baidu Qianfan (international) stores part of its personal data in China. Alibaba’s “International” inference scope excludes mainland China, but its “Global” scope does not. MiniMax (international) discloses only US storage.
- Vendors retain data for four main reasons: abuse monitoring, investigation of flagged content, legal obligations, and stateful features you choose to use. Zero data retention (ZDR) removes only the first. Anthropic still keeps flagged content for up to two years under ZDR.
- DeepSeek, Kimi (international), MiniMax (international) and Baidu Qianfan (international) have no public DPA or processor terms, and Z.ai’s DPA lacks sub-processor and audit terms. Those five will usually not clear a US enterprise privacy review.
This is part 2 of our series on privacy across the AI supply chain, written for teams selling into US and EU markets that are choosing a model API. Part 1 mapped the 10 places a prompt gets stored; this article covers the model vendor layer only. Part one lays out the facts. Part two covers what they mean for a vendor review. For Chinese vendors we looked only at the offerings sold to overseas customers. Terms are current as of September 17, 2026. They change often, so check the linked sources before you sign.
Part one: what the 13 vendors’ terms actually say
What are you doing, legally, when you call an LLM API?
Usually, you are handing personal data to a processor. User data goes into the vendor’s inference service through the prompt, and the vendor processes it on your instructions and sends back a result. Under GDPR that is a controller-processor relationship; under the CCPA it is a business and its service provider. Both laws require a written contract that sets out the purpose and scope of processing, bars the vendor from using the data for anything else, and flows the same restrictions down to sub-processors. GDPR also requires deletion or return of the data when the contract ends.
Once a vendor uses your data for its own purposes, such as training its own models, that processing falls outside the relationship. Under GDPR the role turns on who decides the purposes and means of that processing: a vendor deciding on its own is generally an independent controller, and one deciding jointly with you may be a joint controller. Either way, you need your own legal basis to disclose the data to it. If your business is subject to the CCPA and the vendor uses the data beyond what a service provider contract permits, it likely no longer qualifies as a service provider; the disclosure may then be treated as a “sale” or “sharing,” which consumers have the right to opt out of.
Part 1 condensed these requirements into six questions, which we use throughout:
| # | Question | Relevant law |
|---|---|---|
| 1 | How long are prompts and outputs kept by default, and why? | GDPR, CCPA |
| 2 | Is ZDR available, and which endpoints does it cover? | GDPR, CCPA |
| 3 | Is data used for training by default? Where is the control, and is it opt-in or opt-out? | GDPR, CCPA, FTC Act |
| 4 | Where do inference and storage physically happen? Can you choose a region? | GDPR, DOJ data security rule under EO 14117 |
| 5 | Who can review data manually, and for how long? | GDPR, CCPA |
| 6 | Who is the contracting entity? Is there a public DPA and sub-processor list? | GDPR, CCPA |
Privacy controls: training, ZDR and where inference runs
The same vendor often applies different terms to its consumer, enterprise and API plans, so the figure lists plans separately wherever the handling differs. Among enterprise and API plans, the DeepSeek, Kimi and MiniMax APIs are the ones whose terms clearly allow training or service improvement by default; four others are unclear. Consumer plans and free tiers are the opposite: most train by default, and none offer ZDR.
Enterprise and API plans
How to read the figure: Yes and No mean the terms say so explicitly. Unclear means the terms are silent or contradict each other, None means there is no control to find, and NA means not applicable. The ZDR column shows how you get it (On by default, Self-serve, Approval, Contract) or No. Global means no regional commitment; Global* is Alibaba’s Global inference scope, which does not exclude mainland China.
- OpenAI: ZDR requires approval and covers 12 endpoints. Every region outside the US requires approved ZDR or Modified Abuse Monitoring first. The region you pick for ChatGPT Business covers storage only, and a copy of every prompt is also kept in the US.
- Anthropic: the Claude Team and Enterprise interfaces are not ZDR-eligible, except Claude Code in Enterprise organizations with ZDR enabled. On the API,
inference_geocan currently restrict inference to the US only. - Google: the Gemini Developer API has neither ZDR nor region selection. Since the September documentation update, Google points customers who need ZDR to Vertex AI.
- Mistral and Cohere: Mistral says pay-as-you-go customers “have the right to opt out” of training, and Cohere offers an opt-out toggle; neither states the default, so we mark it Unclear.
- Kimi API: the terms allow Moonshot to use customer content to improve its services, while the help center says API data isn’t used for training. We go with the binding terms and mark it Yes.
- Z.ai: the DPA says API content is not stored. The “will not use to improve services” sentence in the terms of use is published inside literal square brackets, so we mark training Unclear.
- BytePlus: regions are Johor and Dublin, and the documentation says EU traffic may spill over to Asia-Pacific.
A Singapore entity doesn’t mean your data stays out of China. DeepSeek has no overseas entity: calling its API means contracting with a Hangzhou company and having data processed in China. Italy’s Garante restricted DeepSeek in January 2025 because data was “stored in the People’s Republic of China,” and South Korea’s PIPC found in April 2025 that it had sent user input to Beijing Volcano Engine without consent. Both actions targeted DeepSeek’s own app. For API customers, the evidence is DeepSeek’s privacy policy and open platform terms, which also say personal data is collected, processed and stored in the PRC. Alibaba separates region from inference scope: its International scope excludes mainland China, while the Global scope is “dynamically scheduled worldwide” and does not.
For Americans’ data, ownership matters too. The DOJ’s data security rule under Executive Order 14117 looks beyond where data sits to whether the counterparty is a “covered person,” which includes companies organized or headquartered in a country of concern and companies 50% or more owned by one. A Singapore subsidiary of a Chinese parent can qualify. Sending bulk US sensitive personal data to such a vendor for inference may be a restricted transaction that requires CISA’s security requirements.
Anthropic’s Covered Models come with mandatory retention. Per Anthropic’s Covered Models page, Claude Fable 5, Fable 5.1, Mythos 5 and Mythos 5.1 require at least 30 days of retention on the API, Bedrock, Google Cloud and Foundry, for safety purposes. Per its API data retention docs, requests fail with a 400 error if an organization’s retention settings don’t allow it. Anthropic has opened a limited transition that lets eligible customers use Fable 5 and 5.1 with ZDR for internal business applications. Check whether your ZDR agreement accounts for this before you buy.
Consumer plans and free tiers
These plans don’t belong in production. We include them because employees use personal accounts, and you need to know the exposure. None offer ZDR, so the figure has four columns.
ChatGPT personal plans and the Gemini app train by default and let you turn it off in settings. Claude and Grok consumer plans have a control, but the official pages don’t state the default. The Gemini API free tier can’t be opted out of without paying. Qwen Chat (international) processes data in Singapore and mainland China. The Kimi app’s privacy policy names a Beijing entity and storage in China; the page may vary by visitor location, so we mark the entity Unclear.
Why do vendors keep your data?
Four reasons cover most of it, and ZDR and negotiation only reach some of them.
| Purpose | Typical retention | Removed by ZDR? |
|---|---|---|
| Abuse monitoring | 30 to 55 days | Yes |
| Investigating flagged content | Months to years | No |
| Legal obligations and litigation holds | Open-ended | No |
| Stateful features | Per feature, often “until deleted” | Most are unavailable under ZDR |
- Abuse monitoring: OpenAI keeps data “for up to 30 days, unless longer retention is required by law.” Google keeps it for 55 days, and its stated purposes include “legal or regulatory disclosures.” xAI keeps it for 30 days “for auditing purposes in the event of suspected abuse.”
- Flagged content: Anthropic keeps flagged inputs and outputs for up to two years and classifier scores for up to seven, and says the two-year retention applies “even with ZDR or HIPAA arrangements.” BytePlus keeps content caught by its filter for 180 days in Malaysia. OpenAI keeps images flagged by its CSAM scanner for human review even under ZDR.
- Legal obligations: every vendor reserves the right to keep data longer when the law requires it, and none sets a cap.
- Stateful features: OpenAI keeps Conversations, Assistants and vector stores until you delete them. Mistral keeps Agents API data until you close your account. The Gemini Interactions API stores by default (
store=true) for 55 days on the paid tier. Anthropic keeps Batch results for 29 days and code execution containers for 30.
Two more categories sit outside that list. Billing and usage metadata is kept even under ZDR. And “service improvement” is a default purpose for Kimi (international), MiniMax (international) and DeepSeek, while Baidu Qianfan (international) cites “research and statistical analysis.” That is a commercial use, not a safety need, and it’s the first thing to strike in negotiation.
Inference-side caches run on a shorter clock. OpenAI’s in-memory cache expires after 5 to 10 minutes of inactivity, up to an hour, with a separate 24-hour option. Anthropic’s is 5 minutes or 1 hour. DeepSeek’s disk cache is on by default and clears within a few hours to a few days.
Litigation holds override deletion promises. In May 2025, the Southern District of New York ordered OpenAI, in The New York Times v. OpenAI, to preserve output logs that would otherwise have been deleted. Under a stipulated order entered October 9, 2025, the obligation ended as of September 26, 2025, and data already preserved stayed preserved. ChatGPT Enterprise was later carved out. The statement that API customers on ZDR endpoints were unaffected comes from OpenAI, not from the court’s orders. The logic is straightforward: a deletion policy yields to a preservation order, but data that was never stored can’t be preserved.
Do vendors offer different retention terms by industry?
Some do, but most offer the same retention terms with an extra contract attached. Only Anthropic, OpenAI and Google Vertex actually change how data is retained.
Healthcare (HIPAA)
- Anthropic changes the retention model. HIPAA readiness can be turned on self-serve in the Console with the standard business associate agreement (BAA), or negotiated with sales. It applies to the whole organization permanently, and out-of-scope features such as Batch, Files and code execution return a 400 error. Anthropic positions it as an alternative to ZDR: data may be retained, with added safeguards.
- OpenAI layers it on its retention tiers. Its Business Associate and Healthcare Addendum covers designated endpoints, and health data can be processed even when data is retained under Eyes Off or Safety Retention.
- Same terms, plus a contract. Google offers a BAA only through Vertex under the Google Cloud BAA; the Gemini Developer API isn’t on the covered products list. Mistral offers a BAA for qualifying services. xAI reviews BAA requests through a questionnaire. Cohere’s BAA covers custom model projects only, not its SaaS API.
- Explicitly prohibited. Kimi (international), Z.ai and MiniMax (international) prohibit protected health information. For US use, Z.ai also prohibits nonpublic personal information under GLBA and children’s data processed in violation of COPPA.
Government: Anthropic’s Claude for Government is FedRAMP High, and Google offers FedRAMP High and IL4 on Vertex through Assured Workloads. These are compliance boundaries; the retention terms don’t change.
Middle tiers for regulated customers: OpenAI’s Modified Abuse Monitoring keeps prompts out of abuse logs while leaving stateful features available. Eyes Off keeps logs but excludes human review. Safety Retention keeps and reviews content tied to “severe risk activity.” Google exempts Vertex customers under a Google Cloud Master Agreement from prompt logging by default.
Keeping the vendor away from your data: in Cohere’s private, VPC and third-party cloud deployments, “Cohere does not receive any customer inputs.” Mistral offers on-premises deployment and sovereign compute for the public sector. These are deployment models, not retention terms.
If your data is subject to HIPAA, start by removing the three vendors that prohibit PHI, then separate the vendors where a BAA is enough from those that require a different processing configuration. The latter disables some features, so design around the eligible endpoints from the start. None of the 13 vendors publishes retention terms specific to education or financial services; those customers rely on standard enterprise agreements.
Part two: what this means for your vendor review
Each set of facts in part one maps to a step in a vendor review:
- Trains or improves by default, with no control (the DeepSeek, Kimi and MiniMax APIs): the vendor is an independent controller for that processing and likely isn’t a CCPA service provider. For any workload involving personal data, these vendors should be out in the first round unless you can sign an enterprise agreement that changes the default.
- ZDR: attach the vendor’s feature eligibility table to the contract so it’s clear which endpoints, models and features are covered. Flagged content and legal holds are still retained; record both in your risk assessment.
- Inference in China or undisclosed (DeepSeek, Alibaba’s Global scope, Kimi, MiniMax, Baidu Qianfan): you can’t complete a transfer impact assessment. Either get written confirmation of where inference runs or assess the worst case. If the data involves Americans, check ownership under the EO 14117 rule.
- Retention purposes: abuse-monitoring periods and ZDR are negotiable. Flagged content and legal holds aren’t, but you can ask for caps and notice. “Service improvement” should come out.
- Industry terms: for healthcare, remove the three vendors that prohibit PHI, then decide between a BAA and a different processing configuration.
What compliance documents will they give you?
Good terms don’t get you through procurement without the paperwork. US enterprise security teams usually start with the DPA, the sub-processor list, a SOC 2 report and ISO certifications, and the model documentation the EU AI Act requires.
How to read the figure: “Can terminate” means whether you can end the service if an objection to a new sub-processor isn’t resolved. Partial means the list omits the service itself. Platform means the certification covers the cloud platform but doesn’t name the model service. Excluded means the Gemini API is outside Google’s ISO 42001 scope. Custom means a training data disclosure that doesn’t follow the EU template. Thin means a DPA without sub-processor, audit or breach notification terms.
- OpenAI’s sub-processor list includes outsourced content moderation in the Philippines, and samples of flagged content are shared with those providers.
- Anthropic publishes its model documentation, training data summaries and ISO 42001 certificate. Mistral has the most complete AI Act materials, with public training summaries and downstream technical documentation for 41 models.
- Google’s ISO 42001 scope includes Gemini Enterprise Agent Platform but not the Gemini API. None of BytePlus’s 16 certifications names ModelArk.
- xAI’s trust center is a request form, open only to customers under NDA. Cohere is the only vendor that requires an NDA just to see its DPA.
- Alibaba’s DPA lets it refuse audit reports to competitors and provide only a “summary copy.”
- BytePlus has the most complete DPA among the Chinese vendors, with EU SCCs and UK, Swiss, Brazilian and Saudi addenda. But ModelArk isn’t on its sub-processor list, and the only remedy for an objection is “reasonable efforts” to change the service.
Which gaps will block procurement? Against typical US enterprise review standards, the gaps fall into three tiers:
- Fails privacy review: no DPA or equivalent processor terms, or no sub-processor list. GDPR requires a written contract with the Article 28 terms between controller and processor, and the CCPA requires a service provider contract with specific restrictions. The document doesn’t have to be called a DPA, but if the standard terms don’t contain those provisions, the required contract terms are missing. GDPR doesn’t require a published sub-processor list, but without one you can’t verify what you’ve authorized or meaningfully object to changes, and most US enterprise procurement standards treat that as a gap. DeepSeek, Kimi (international), MiniMax (international) and Baidu Qianfan (international) fall here. Z.ai’s DPA lacks sub-processor, audit and breach notification terms and won’t pass a check against GDPR’s required clauses either.
- Fails security review: no SOC 2 report. A report available under NDA counts. DeepSeek, Kimi (international) and Z.ai have none, and we couldn’t find one for MiniMax (international) or Baidu Qianfan (international).
- Worth recording, not blocking: ISO 42001 isn’t a hard requirement yet, but record the certification scope so a platform-level certificate isn’t logged as covering the model service. If you deploy in the EU and the vendor has no AI Act training content summary, you’ll have to fill that gap in your own documentation. Reports behind an NDA or request form only add time.
Three more things worth noting.
Check the ISO 42001 scope. OpenAI, Anthropic, Google and Cohere hold it at the company level, but Google’s scope excludes the Gemini API, and none of the Chinese vendors’ overseas offerings holds it. Fireworks AI and GitHub Copilot have it too, so it isn’t out of reach; Mistral, the vendor most directly subject to the AI Act, doesn’t. Record the scope, not just “certified.”
Sub-processor objection periods range from 10 to 90 days. Google gives 90 days with a termination right. OpenAI gives 30. Anthropic gives 15. xAI gives 15 with a termination right. BytePlus gives 15 plus 15, with only a “reasonable efforts” remedy. Mistral gives 10. Alibaba gives 10 business days with a right to terminate the affected service. Cohere publishes its list but not its objection period.
A vendor’s own documents often disagree. Z.ai’s no-improvement clause is published in square brackets, like an unresolved drafting note. BytePlus’s 2022 terms still grant a “perpetual, irrevocable” data license that its new customer agreement narrows to technical data, and its data processing page lists Indonesia while its regions page lists only Johor and Dublin. xAI’s enterprise terms require personal data to go only through ZDR endpoints, while its developer docs say “for most customers, we do not recommend enabling ZDR.” Your review record should name the document, date and plan you relied on.
Which questionnaires have they already completed?
Very few, and two of the four common questionnaires don’t ask about AI at all.
CSA CAIQ and AI-CAIQ. The Cloud Security Alliance released version 1.1 of its AI Controls Matrix in June 2026, with 247 control objectives; the companion AI-CAIQ treats LLM API vendors as “model providers.” As of September 2026, the CSA STAR Registry held 2,737 records; searching by each vendor’s legal entity name, none of the 13 vendors has submitted an AI-CAIQ. OpenAI has a CAIQ from April 2024 and xAI a CAIQ v4.0.3 from January 2026. Anthropic is listed only with ISO 42001; under CSA’s own definition, STAR for AI Level 2 requires both ISO 42001 and a validated AI-CAIQ. Alibaba, ByteDance (as Beijing Volcano Engine Technology) and Baidu have records, none of them STAR for AI. Mistral, Cohere, DeepSeek, Moonshot, MiniMax and Zhipu have none. Search the registry by legal entity name, or you’ll miss ByteDance and Alibaba.
HECVAT. Version 4.1.6 of the questionnaire most US universities require has 31 AI questions and 8 privacy questions. EDUCAUSE retired its public Community Broker Index in July 2025 and says it doesn’t collect completed HECVATs, so you have to ask each vendor. Anthropic lists HECVAT v4.1.6 in its trust center. OpenAI lists “HECVAT Full” and “HECVAT Lite,” names from version 3; version 4 merged them into a single questionnaire.
SIG. The Shared Assessments questionnaire common in financial services and large enterprises costs $7,000 a year for a corporate license. Anthropic’s trust center lists SIG Lite and a full CAIQ, available on request rather than filed with the STAR Registry.
VSA. The Vendor Security Alliance’s two questionnaires are free, but the “current” version was last updated in February 2022 and doesn’t mention AI, machine learning, LLMs or generative models anywhere.
AI-specific questionnaires are almost never pre-completed. If you want answers, you’ll have to send the questions yourself.
What can you negotiate, and what can’t you?
Checked against GDPR’s required DPA terms, the gaps cluster in four areas.
Regions and cross-border transfers. You can lock a region in configuration with OpenAI (non-US regions require approved ZDR or Modified Abuse Monitoring), Mistral’s regional endpoints, Alibaba’s geographically bounded inference scopes, xAI’s US endpoint (grok-4.6 only) and Anthropic’s inference_geo (US only). The Google Gemini Developer API and Cohere offer no region choice. BytePlus has a Dublin region but says EU requests may spill over to Asia-Pacific. If your privacy notice tells EU users their data stays in the EU, verify it endpoint by endpoint, not vendor by vendor.
Human access. All six US, EU and Canadian vendors keep a human review path. OpenAI shares samples of flagged content with sub-processors, and its Eyes Off tier excludes human review. Anthropic says no employee can read retained Covered Models conversations by default and that each access is logged in a tamper-evident log. The Chinese vendors’ overseas offerings generally keep automated content screening.
Sub-processors. OpenAI, Anthropic, Mistral, xAI, Alibaba and BytePlus publish a list and an objection period. Google has the contractual terms but no Gemini-specific list. Cohere publishes the list but not the objection period. The other five have no list.
Deletion at termination. Mistral and Anthropic commit to 30 days and BytePlus to 180. OpenAI and Google give no number. None of the 13 discloses its backup purge cycle.
Before you sign, negotiate at least these five points:
- Retention and ZDR eligibility by endpoint. Attach the feature eligibility table to the contract rather than pointing to a web page that can change.
- A cap on flagged-content retention, including who can see it and whether you’re notified.
- Inference location and storage region as separate commitments. A storage commitment alone doesn’t stop inference from crossing borders, and “global” or “international” scopes should come with a country list.
- The sub-processor list as a contract exhibit, with an objection period and a remedy. For vendors without a list, draft this from scratch.
- Deletion timing at termination, backup purge cycles and notice of litigation holds. xAI’s statement that customers on ZDR are “solely responsible for preserving any copies” is a useful model for drawing the line.
Some terms won’t move. Grounding logs for Search and Maps on the Gemini Developer API are kept for 30 days and can’t be turned off (3 days on Vertex). Cohere won’t show its DPA without an NDA. Kimi (international), MiniMax (international), Baidu Qianfan (international) and DeepSeek have no public DPA, so you’d be negotiating from a blank page. Anthropic’s 30-day retention for Covered Models started as a condition of access, but the ZDR transition suggests it’s softening; ask again at renewal.
Ten questions to put in your review
We picked ten questions from HECVAT 4.1.6 and the AI-CAIQ. Question IDs follow the official EDUCAUSE workbook; an asterisk marks questions HECVAT flags as critical.
- If sensitive data gets into your AI model, can it be removed on request? (AISC-01*) Deleting a conversation and removing data from a model are different things.
- Are user inputs used to influence your AI models? (AISC-02*, including fine-tuning, training and personalization) This decides whether the vendor is a processor or an independent controller.
- Is institutional data retained during AI processing? (DPAI-02*) Ask for retention periods, purposes and differences by feature.
- Is institutional data processed by third parties or sub-processors that also use AI? (DPAI-04) Especially relevant for aggregator gateways and cloud-hosted models.
- Do you have agreements with sub-processors covering data protection and AI use? (DPAI-03*) GDPR requires obligations to flow down.
- Can AI features be disabled per tenant or per user? (AIGN-02*) Employee use and purpose limitation both depend on it.
- Are there tamper-evident logs with user, time and action, available for audit? (AISC-03*) Audit rights are only as good as the logs behind them.
- Is training data separated from customer data, and can customers opt out of model improvement? (AIML-01*) Separation is what makes a no-training commitment real.
- Do actions taken by LLM features or plugins require human approval? (AILM-03*) A key control for agents and tool use.
- Is institutional data processed by any shared AI service? (DPAI-06) Multi-tenancy is the common root of cache isolation, abuse review and sub-processor questions.
Questions 6 and 9 look like questions for the vendor, but they’re really about you. Once you build an LLM API into a product you sell, you’re the one answering the HECVAT, and whether you can say “yes” depends on whether your model vendor and gateway give you the control.
What your compliance team should do next
First, add every LLM API in use to your record of processing activities, with at least the contracting entity, inference location, storage location, retention periods and purposes, training default and ZDR status. For Chinese vendors’ overseas offerings, verify inference location separately and treat “global” or “international” without a country list as undisclosed. If the data includes Americans’ sensitive personal data, check ownership as well.
Second, update the processor list and cross-border transfer section of your privacy notice. For EU users, Cohere can only be described as processing in the US, the Gemini Developer API as processing in “any country” where Google has facilities, and DeepSeek as processing in China.
Third, run a data protection impact assessment. The EDPB’s Opinion 28/2024 expects deployers to assess whether a model was developed with lawfully processed personal data. France’s CNIL is blunter: when you use an external API, “control of the system lies almost entirely with the provider,” and the relationship has to be governed by a data processing agreement.
Fourth, get personal accounts out of production workflows. ChatGPT and Gemini consumer plans train by default, and the Claude consumer plan has a control whose default isn’t documented, while the API and enterprise plans from the same vendors don’t train by default. An employee pasting business data into a personal account puts it on the training-enabled side.
Finally, put terms reviews on a schedule. In 2026 alone these 13 vendors changed their terms more than a dozen times: Google’s Gemini API ZDR page and xAI’s security FAQ both changed in mid-September, xAI’s legal pages now name SpaceXAI LLC, and BytePlus’s new customer agreement takes effect September 30. A one-time review at signing isn’t enough; terms reviews belong on the same recurring calendar as key rotation.
Frequently asked questions
Do LLM APIs train on my data by default?
It depends on the vendor and plan. The OpenAI, Anthropic, paid Google Gemini, xAI, Alibaba (international) and BytePlus APIs don’t train by default. The DeepSeek, Kimi (international) and MiniMax (international) APIs use inputs and outputs to train or improve their services by default. Mistral pay-as-you-go and Cohere offer an opt-out without stating the default, so check the admin console. Nearly every vendor’s consumer plans and free tiers train by default, which is why personal accounts don’t belong in production.
Does zero data retention mean the vendor stores nothing?
No. ZDR usually removes only abuse-monitoring logs, and coverage depends on the vendor’s feature eligibility table. Flagged content (up to two years at Anthropic), billing metadata and data from stateful features you use are still retained, and OpenAI keeps CSAM-flagged images even under ZDR. The Gemini Developer API has no ZDR at all; you need Vertex AI.
If I use a Chinese vendor’s international offering, does my data stay out of China?
Not necessarily. A Singapore contracting entity tells you where the contract is signed, not where data goes. DeepSeek processes and stores data in China. Baidu Qianfan (international) stores part of its personal data in China. Alibaba’s Global inference scope doesn’t exclude mainland China, so choose a bounded scope such as International or US. MiniMax (international) and the Kimi API disclose storage location but not where inference runs. For Americans’ data, also check ownership: an overseas subsidiary controlled by a Chinese company may be a covered person under the EO 14117 rule.
Which LLM API vendors will sign a HIPAA business associate agreement?
OpenAI, Anthropic, Mistral and xAI will. Google signs only through Vertex AI; the Gemini Developer API isn’t on its covered products list. Cohere’s BAA covers custom model projects only. Kimi (international), Z.ai and MiniMax (international) prohibit protected health information, and DeepSeek says its service isn’t intended for sensitive data such as health information.
What’s the compliance problem when employees use personal ChatGPT or Claude accounts with customer data?
You’ve disclosed customer personal information to a third party with no data processing agreement. Consumer terms aren’t a GDPR processing agreement and don’t meet the CCPA’s service provider contract requirements, so the vendor processes the data under its own privacy policy and you have no instruction, deletion or audit rights. The fix has two parts: move business use to API or enterprise accounts, and give employees a written AI use policy that spells out what’s allowed.
Appendix: sources
US, EU and Canadian vendors: OpenAI data controls, prompt caching, DPA v.010126, sub-processor list, response to NYT data demands, trust portal, business data, enterprise privacy, ChatGPT data residency, ChatGPT Business data residency, ChatGPT data controls; Anthropic API and data retention, data residency, Covered Models, Covered Models retention, commercial data retention, commercial training policy, Development Partner Program, Claude for Government, ZDR scope, Enterprise custom retention, consumer training and retention, DPA, trust center and sub-processors; Google Gemini API terms, usage policies, ZDR, Interactions API, processor terms, sub-processors, ISO 42001 scope, HIPAA covered products, Vertex abuse monitoring, Vertex ZDR, Gemini for Government, Vertex data residency, Gemini Apps privacy hub, Workspace generative AI privacy hub, Workspace data regions coverage; Mistral privacy policy, DPA, ZDR, training opt-out, privacy controls, data hosting, regional inference, sub-processors, trust center resources, AI governance model list; xAI enterprise terms, enterprise FAQ, consumer FAQ, Grok Business announcement, DPA, security FAQ, regional endpoints, sub-processors, trust center; Cohere enterprise data commitments, data usage policy, privacy policy, trust center, sub-processors, BAA for custom models.
Chinese vendors’ overseas offerings: DeepSeek privacy policy, app terms of use, open platform terms, context caching; Alibaba Cloud Model Studio (international) regions and inference scopes, privacy notice, product terms, DPA, membership agreement, inference scope details; Qwen Chat privacy policy; Kimi (international) terms of service, privacy policy, API data security, Kimi app privacy policy; Z.ai privacy policy with DPA, terms of use, Z.ai Chat privacy policy, Z.ai Chat terms; BytePlus ModelArk data processing, regions, content pre-filter, service-specific terms, customer agreement, DPA, sub-processors, certifications, data authorization agreement; MiniMax (international) privacy policy, terms of service, MiniMax Agent privacy policy, MiniMax Agent terms; Baidu Qianfan (international) user agreement, privacy policy.
Law and regulators: GDPR; California Consumer Privacy Act (Cal. Civ. Code title 1.81.5); CPPA regulations; EDPB Opinion 28/2024; CNIL Q&A on generative AI; DOJ Data Security Program; Garante decision 33/2025; Korea PIPC notice, April 24, 2025; In re OpenAI Copyright Infringement Litigation, No. 1:25-md-03143 (S.D.N.Y.): preservation order, May 13, 2025, stipulated order ending preservation, October 9, 2025.
Compliance documents and questionnaires: CSA AI Controls Matrix v1.1 and AI-CAIQ, STAR for AI levels, STAR Registry; EDUCAUSE HECVAT 4.1.6; Shared Assessments SIG; Vendor Security Alliance; Mistral compliance hub; ByteDance Seed transparency; Alibaba Qwen and Wan training data disclosure; Fireworks trust center, GitHub Copilot trust center; European Commission training content summary template.
This article is an engineering and compliance reading of vendor terms, not legal advice. Part 1 of the series: AI Data Compliance Map.
