AI & Privacy
Most AI is training on you.
Nearly every consumer chatbot now feeds your conversations back into its models by default — the opt-out is buried, and anything you already typed is gone. Here's what the major companies actually do with your data, judged on privacy alone, plus the options that don't work this way.
A ground rule: we judge AI on privacy alone — what each company does with what you type. An institutional deal doesn’t count on the open market (Pitt’s enterprise agreement with Anthropic is real and worth using, but that’s a campus thing). Which assistant falls where is the ranking.
Set the deals aside and most consumer AI lands in the same place. Paying $20 a month doesn't fix it: personal Plus and Pro plans usually still train on you. The tiers that don't are enterprise/education accounts (contractually excluded) and models you run yourself (nothing leaves your device).
The test
How to judge any AI on privacy
The questions that actually matter — the same ones we apply to everything else. Run any tool, new or old, through these.
- Trained on your chats by default? On consumer plans, usually yes. That's the first thing to check and change.
- Is there a real opt-out? And is it honest, or a buried toggle designed so you never find it?
- How long is data kept? Retention runs from ~30 days to several years — even after you hit delete.
- Do humans read it? Many services let staff or contractors review conversations for "safety."
- Is it wired into an ad or data empire? A chatbot owned by an ad company feeds the same profile as your searches and email.
- What jurisdiction, and can you self-host? Where the servers sit decides who can compel the data — and open-weight models can skip the cloud entirely.
Same standard we hold everywhere — see how we evaluate what we publish.
A note on Gmail. Google’s “Smart Features” — which power autocomplete, category sorting, suggested replies, and Gemini’s ability to summarize your inbox — scan your email content by algorithm. This isn’t new: it has worked this way since roughly 2014. Google says Gmail content is not used to train its general-purpose Gemini model; the scanning is for personalizing your own experience. In November 2025 Google added an opt-in Gemini Deep Research mode that can pull from your Gmail, Drive, and Calendar, but that requires a separate choice. The honest summary: Gmail scans every message algorithmically by default; no one at Google reads your email; your content shapes your profile but (per Google) doesn’t train the shared model. If any of that bothers you, Smart Features is a toggle — Settings → See all settings → General → Smart features and personalization → off. The trade-off is losing suggestions and Gemini sidebar access.
The field
Where does your prompt actually go?
One question splits every AI apart. If what you type never leaves your device, privacy is simply a fact. The moment it reaches a company, you’re left with a promise — or, at best, a promise you can check. And once it’s with a company, the law protects it far less than what stays on your device — that’s the third-party doctrine.
On-device tools like Ollama and LM Studio keep everything local. A few services — Apple, Tinfoil, Brave’s TEE — let you cryptographically check the company can’t read your prompt. Everything else, ChatGPT and Claude and Gemini included, is a promise whose worth depends on the law where the servers sit — we put all 31 through the same test.
The other kind of training
If you make things, you can fight back
Everything above is about your chats. But the same models train on scraped images — artists keep finding their work in training sets like LAION-5B with no consent, credit, or pay, then watch it used to mimic their style on demand. Two free tools from the University of Chicago’s SAND Lab let creators push back.
Defensive · cloak your style
Glaze
Adds changes invisible to your eye that make AI models read your piece as a completely different style, so a model fine-tuned to copy you learns the wrong thing. It runs on your own machine and sends nothing back. WebGlaze is a free, invite-only web version for phones or computers without a GPU.
Offensive · poison the scrape
Nightshade
Alters an image so a model that scrapes it learns the wrong associations — a poisoned dog reads as a cat. The goal isn’t to break models; it’s to make training on unlicensed work costly enough that paying to license it becomes the easier path.
Neither is permanent — it’s an arms race, and the team updates both as new attacks appear. Think of them as a stopgap while the law catches up, not a guarantee. Sources: SAND Lab — What Is Glaze · MIT Technology Review on Nightshade.
Get involved
It’s free, open to everyone at Pitt, and joining takes about a minute.
No dues, no experience needed. Come to a meeting, or leave your name and we’ll tell you when the next one is.