⚡ Key takeaways
- Cloud AI privacy depends on the tier and settings you use, not on the brand. Business and API plans usually default to no training; consumer chat often does not.
- Deleting a chat does not remove every copy. Providers document abuse-monitoring logs, safety reviews and flagged-content retention.
- Only local inference keeps prompts on your own machine. A local app that calls a cloud model does not count.
You paste a contract, a medical email or some source code into an AI chat, and one question follows: who else can see this? The honest answer in the local vs cloud AI debate is that it depends on the product, the plan and one or two settings most people never open. This guide walks through what the major providers say they keep, who may review it, and what really stays on your device when you run a model yourself.
Local vs cloud AI: the basic difference
With a cloud model, your prompt travels to the provider’s servers, gets processed there, and the provider decides how long it keeps the input and output. With a local model, the same work happens on your own hardware. Ollama, a local runner, states in its privacy policy that it does not collect, store, transmit or have access to prompts, responses or content processed locally. Its own wording is short: “Your data stays on your machine.”
Cloud services are not one thing, either. The same company often runs a consumer product and a business product with opposite defaults. That split is where most misunderstandings come from.
What the big cloud providers keep
| Provider | Consumer chat default | Business or API default |
|---|---|---|
| OpenAI | May train on your content; opt-out available | Not used for training |
| Anthropic | Training only if you allow it | Retained data not used for training without permission |
| Human review possible; auto-delete on a timer | Depends on tier: paid and Workspace are stricter | |
| Microsoft | Not covered here | Prompts and responses not used to train foundation LLMs |
OpenAI
OpenAI’s platform documentation says data sent to the API has not been used to train or improve its models since March 1, 2023, unless you opt in. Business tiers follow the same rule by default:
- ChatGPT Enterprise
- ChatGPT Business
- ChatGPT Edu
- ChatGPT for Healthcare
- ChatGPT for Teachers
- The API platform
OpenAI’s enterprise privacy page says it retains API inputs and outputs for up to 30 days, and the platform docs describe those abuse-monitoring logs the same way.
Consumer ChatGPT and Codex work differently. OpenAI’s help article says it may use your content to train its models, and you can opt out. There is a catch: after you opt out, the full conversation tied to any thumbs-up or thumbs-down feedback may still be used. Temporary chats are not used to improve models while they stay temporary, but a saved one follows your account settings.
Zero Data Retention (ZDR) is granted by approval, for eligible endpoints only. Some endpoints that store application state, such as conversations and assistants, are not ZDR-eligible. Images flagged for potential child sexual abuse material are still retained for manual review and reporting, even under ZDR. Zero Data Retention is not limited to OpenAI. If you use another agent framework, see our guide to enabling zero data retention in Hermes Agent with Command Code for how it works outside the OpenAI API.
Anthropic
For Claude’s consumer plans, Anthropic’s Privacy Center says deleted chats leave back-end storage within 30 days. If you allow your data to be used for model training, new and resumed chats are kept in de-identified form for up to 5 years. Incognito chats are not used to improve Claude, even with Model Improvement enabled.
Flagged content is the exception to deletion. Anthropic keeps flagged inputs and outputs for up to 2 years and trust-and-safety classification scores for up to 7 years. On the API side, the API retention docs say inputs and outputs are not retained by default, except for Covered Models, which require 30-day retention. In Anthropic’s words: “Retained data is never used for model training without your express permission.”
With Google, the setting you pick changes the outcome. In the consumer Gemini Apps, chats auto-delete after 18 months by default, and you can change that to 3 months, 36 months or indefinite. Chats that human reviewers have looked at are not deleted when you delete your activity; Google keeps them for up to 3 years. Google’s own warning is blunt: do not enter confidential information you would not want a reviewer to see.
Turning Keep Activity off helps less than it sounds. Chats are still retained with your account for 72 hours, and human review remains active. Temporary chats are not used to train Google’s AI models, but they may still be reviewed for safety purposes.
The developer and business products are stricter, with one trap. The Gemini API is meant for professional or business use, not consumers, and its terms say that on the unpaid tier human reviewers may read, annotate and process your input and output. The paid tier does not use prompts to improve Google’s products. Vertex AI says Google will not train or fine-tune models on your data without prior permission. Reaching zero data retention there takes care: Grounding with Google Search is optional, but if you use it, prompts and outputs are stored for up to 3 days and that storage cannot be turned off. Google’s Workspace privacy hub, updated October 6, 2026, says Gemini content in qualifying Workspace editions is not human reviewed and not used for training outside your domain without permission, with prompts and responses kept from 90 days to indefinitely, as set by admins.
Microsoft
Microsoft’s documentation for Microsoft Copilot, the work product formerly called Microsoft 365 Copilot, says prompts, responses and data accessed through Microsoft Graph are not used to train foundation LLMs. That data is processed and stored under your organization’s Microsoft 365 contractual commitments and encrypted at rest. This page covers the business product, so do not assume the same terms for any other Copilot you use.
Deleting is not erasing
With OpenAI, Anthropic and Google, “delete” is not the same as “gone”. OpenAI keeps abuse-monitoring logs and flagged imagery, Anthropic keeps flagged content and safety scores, and Google keeps reviewed chats. Treat deletion as removing the chat from your view, then read the exceptions before you paste anything sensitive.
What stays private when you run a model locally
Local inference means the model weights and your prompts both live on your hardware. Ollama exposes a REST API on localhost port 11434, and it builds on llama.cpp, an MIT-licensed C/C++ engine for running LLMs on a wide range of hardware. Ollama’s privacy policy adds: “Ollama runs locally. We don’t see your prompts or data when you run locally.” For the tools that sit around a local model, see our guide to the local AI stack.
The word “local” can mislead you, though. Ollama also offers cloud models that run on remote hardware. Its policy says they process prompts and responses transiently and never train on them, but those prompts do leave your machine. Check which model you actually selected before you assume a prompt stayed home.
The hybrid case: Apple Intelligence
Apple splits the work. Apple Intelligence is on-device first, with a roughly 3-billion-parameter on-device model in its 2024 generation, and complex requests go to Private Cloud Compute (PCC). Apple’s security team says PCC data must not be retained after the response, and that user data must never be available to anyone other than the user, “not even to Apple staff”. In a June 8, 2026 post, Apple described extending PCC to infrastructure on Google Cloud, using NVIDIA GPUs, Intel CPUs with TDX and Google’s Titan chip. On training, Apple’s machine learning research team says it does not use users’ private personal data or interactions to train its foundation models.
How to choose
- Highly sensitive material (client files, health records, credentials): run a local model, and confirm it is not a cloud-hosted variant.
- Work use with a team: pick a business or API tier, then read its retention and review terms rather than the marketing page.
- Casual use: a consumer chat is fine for non-sensitive prompts. Check the training toggle, use temporary or incognito modes, and assume safety review is still possible.
Conclusion
If a prompt must never leave your machine, local inference is the only setup that gives you that by design. Before your next sensitive paste, open the privacy page for the product you actually use and check three things: the training default, the retention window and who can review flagged content.
Liked this? There's more where it came from.
Get our digest — articles worth your time, no spam, unsubscribe in one click.
Subscribe to Factnetize →