There are two unhelpful positions on AI and privacy. One is that it does not matter because nobody cares about your data. The other is that everything you type is being read by someone.
The reality is specific and knowable, and it varies enormously depending on which product tier you are using.
The path a prompt takes #
When you type into a cloud AI product, your text leaves your device over an encrypted connection and hits the provider’s servers, where it is processed by the model. The model itself does not remember anything between requests. Each request is stateless, and any memory you experience is the application layer resending previous messages back into the context window.
It is usually logged. Providers retain requests for a period, typically thirty days, for abuse monitoring, debugging, and legal compliance. This is standard for essentially all API services and is disclosed.
Employees can access it under specific conditions, typically for abuse investigation or debugging, under access controls and audit logging. Not casually, and not never.
And it may or may not be used for training, which is the part that differs most between tiers.
The tier distinction that actually matters #
This is the most important thing to understand, and most people do not know it.
Consumer free and paid chat products frequently use your conversations for model improvement by default. This is disclosed in the terms and there is usually a setting to turn it off. The setting is often not on by default, and it is often several menus deep.
Business, enterprise, and API tiers typically do not train on your data by default, and this is generally a contractual commitment rather than a setting. The reason is straightforward: companies will not adopt these tools otherwise, and enterprise customers have lawyers.
So the same company can offer two products where one trains on your input and one contractually does not. If you are pasting anything sensitive into a consumer chat tier, that is the thing to check first.
Go into settings on whatever you use, find the training or data controls, and turn training off. It takes two minutes. Do it now rather than resolving to do it later.
What “training on your data” actually means #
The phrase produces more fear than it warrants, and the reality is still worth taking seriously.
Your conversation does not get stored in the model as a retrievable record. Training adjusts billions of weights by tiny amounts based on enormous quantities of text. A single conversation among billions has a vanishingly small influence on the resulting parameters, and there is no lookup table anyone can query to retrieve it.
But memorization is real. Models demonstrably can reproduce sequences from their training data, particularly text that appeared many times, or that is unusual and distinctive. Research has repeatedly extracted verbatim training data from models. The risk concentrates on things that are both rare and repeated: an API key that appears in several of your conversations is a meaningfully different risk profile from a paragraph about your weekend.
So training on your data is a real but modest risk for ordinary content, and a genuine risk for credentials, unique identifiers, and confidential material that you paste repeatedly.
The other exposures people forget #
Training gets the attention. These are more likely to actually affect you.
Breaches. Any stored data can leak, and chat logs are stored data. Providers have had incidents, including one well-publicized case where users briefly saw other users’ conversation titles. Nothing about AI companies makes them immune to the security problems every other company has.
Legal process. Stored conversations are discoverable and subpoenable. If a provider retains logs, those logs can be compelled. This has come up in actual litigation and it will come up more.
Third-party wrappers. An enormous number of AI apps are thin layers over someone else’s API. Your data goes to the wrapper company and to the underlying provider, under two different privacy policies, with a small startup’s security practices in the middle. This is the highest-risk category and the one people scrutinize least.
Browser extensions and integrations. An AI extension with permission to read page content can read everything on every page, including your email and your banking session. The permission model here is extremely coarse.
Your own employer. Enterprise AI deployments frequently log usage, and admins may have access to prompts. Using the company AI tool for a personal question is not private in the way people assume.
Policy changes. Privacy policies are amendable. A commitment made today by a company that gets acquired in three years is worth what the acquirer decides it is worth.
What not to paste #
Regardless of tier, a short list of things that should not go into a cloud AI product:
- Passwords, API keys, tokens, or private keys. If you already did, rotate them.
- Full social security numbers, government IDs, financial account numbers.
- Patient health information, unless under a business associate agreement with the provider.
- Client confidential material, if you have a professional duty of confidentiality. Lawyers, therapists, doctors, and accountants have obligations that are not satisfied by “the terms say they do not train on it.”
- Unreleased business material where a leak would matter.
- Other people’s personal information. They did not consent to this and you are making a decision on their behalf.
That list is not paranoid. It is roughly the same list you would apply to pasting something into any third-party web service, which is what this is.
The architectural answer #
Everything above is about policies: what a company says it will do with your data. Policies are commitments, and commitments can change, expire, get reinterpreted, or be overridden by a court.
There is a category of answer that does not depend on trust at all. If the model runs on your device and the app has no network path for your data, the data cannot leave. That is not a promise, it is a property of the system.
This is the reason I build things this way. Personal LLM runs language models entirely on the phone. No account, no server, no subscription, nothing to breach and nothing to subpoena, because the conversations exist in exactly one place and it is your device. The build write-up covers the engineering.
The case is sharper for audio. Private Transcribe runs Whisper locally, so a recording of a medical appointment, an interview with a source, or a confidential call is transcribed without the audio ever being uploaded. Cloud transcription services upload every recording by definition, and audio is more revealing than most text people worry about. That build story is here, and the technical background is in how speech recognition works.
The limitation, stated plainly: on-device models are less capable than frontier cloud models. For hard reasoning and complex work, the cloud model is better and it is not close. Local does not replace cloud. There is a large and growing set of tasks where local is good enough, and for those tasks the privacy is free. On-device AI explained works through where that line currently sits.
A practical checklist #
If you do nothing else:
- Turn off training in your settings on every AI product you use. Two minutes.
- Know which tier you are on, since consumer and business tiers have materially different data terms.
- Do not paste secrets: credentials, IDs, financial and health data.
- Be skeptical of wrapper apps. Two companies now hold your data instead of one, and one of them is small.
- Audit your browser extensions and look at what permissions they actually hold.
- Use local models for genuinely sensitive material. This is the only option that does not require trusting anyone.
- Delete old conversations if the product supports it. Less stored data is less exposure.
The proportionate view #
For most people asking most questions, cloud AI is a reasonable thing to use and the risk is comparable to using any other cloud service, which is to say real but ordinary. Email, cloud storage, and messaging apps all have the same structure and nobody agonizes over them.
What is genuinely different is that people paste far more revealing things into an AI chat than into a search box, because the interface feels conversational and private. That is a real behavioral shift and it is worth being conscious of.
Set the toggle, know your tier, keep secrets out, and use a local model for the things that actually matter. That covers most of it.