KYC Drift: Hosted AI — Identity & Data-Handling Facts
Ask a simpler question than “is this AI anonymous?” — because that question has no honest yes/no answer. Ask instead: which identity anchors does each provider’s own paperwork attach to your use, and what does it say happens to what you type? This dataset records exactly that, for eight hosted-AI services, from their own terms, privacy policies, and help pages — each value linked to an archived copy of the document that states it.
The matrix#
| Dimension | ClaudeAnthropic | DeepSeekDeepSeek | GeminiGoogle | Le ChatMistral AI | ChatGPTOpenAI | OpenRouterOpenRouter | PerplexityPerplexity | VeniceVenice |
|---|---|---|---|---|---|---|---|---|
| Identity at signup & use What identity the provider requires/collects at each touchpoint. | ||||||||
| Account required | required | not assessed | required | conditional | not assessed | required | not assessed | optional |
| Signup identifier(s) | not assessed | email or phone | required | not assessed | one of: email / phone / SSO | not assessed | required | not assessed |
| Phone at signup | required ⚠ ↻ | email or phone | not assessed | collected if you provide it | one of: email / phone / SSO | not assessed | not assessed | not assessed |
| Email at signup | conditional | email or phone | not assessed | collected | one of: email / phone / SSO | conditional | not assessed | not assessed |
| Login methods (SSO/wallet) | Google SSO or email magic-link | not assessed | not assessed | not assessed | not assessed | not assessed | not assessed | email, social login, or Web3 wallet |
| Minimum age | 18+ | not assessed | not assessed | 13+ | not assessed | not assessed | not assessed | not assessed |
| Later ID/phone check | conditional | not assessed | not assessed | not assessed | not assessed | not assessed | not assessed | not assessed |
| Payment & billing identity | not assessed | not assessed | not assessed | not assessed | required | accepted | not assessed | not assessed |
| Data handling What happens to your prompts/data (stored, retained, trained on, reviewed, shared). | ||||||||
| Prompts sent downstream | not assessed | not assessed | not assessed | not assessed | not assessed | sent to downstream provider | not assessed | zero data retention (imposed downstream) |
| Inputs used to train | on by default; opt-out available | on by default; no stated opt-out | not assessed | on by default; opt-out available | on by default; opt-out available | varies by downstream provider — selectable | email content excludednot disclosed | not assessed |
| Human review of chats | not assessed | not assessed | yes | not disclosed | not assessed | not assessed | not assessed | not assessed |
| Retention | not assessed | not assessed | up to 3 years | until you delete it | not assessed | not assessed | not assessed | not retained |
| Deletion & identifier reuse | available | available | not assessed | available | availablereusable after deletion | not assessed | not assessed | not assessed |
| Privacy lever (temp chat/ZDR/opt-out) | not assessed | not assessed | available | available | not assessed | not assessed | not assessed | not assessed |
| Data-controller jurisdiction | not assessed | China (CN) | not assessed | FR | not assessed | not assessed | not assessed | not assessed |
| Network layer (separate) OUT OF the provider-identity decoupling scope. Recorded as fact only. IP/device/analytics are network-layer, not provider-account identity. | ||||||||
| Third-party site analytics | not assessed | not assessed | not assessed | not assessed | not assessed | present | not assessed | not assessed |
| Metadata collected (IP etc.) | not assessed | not assessed | collected | not assessed | not assessed | not assessed | not assessed | collected |
Columns are alphabetical. · Every value is what the provider's own document states — not verified enforcement. · not assessed = not yet extracted from primary sources — an empty dimension is NOT evidence that no requirement exists · ⚠ = conflicting sources · ↻ = changed over time · No scores, no rankings, no recommendations.
How to read this#
- Every value is a provider statement. We record what the provider’s published documents say — we do not test enforcement, and a policy saying “we collect X” is not the same as a signup screen refusing to proceed without X. Where those differ, we record them separately.
- “Not assessed” means exactly that. We have not yet extracted that dimension from primary sources. It is not evidence that no requirement exists — treating silence as “no requirement” is the single most common way comparison tables mislead, so this table never leaves a cell blank.
- ⚠ marks a live conflict. When sources disagree (official help says one thing, a third-party report claims another), we show the official statement as the value and record the disagreement on the service page — we do not silently pick a side.
- ↻ marks drift. Requirements change. Where we have evidence of an earlier state, the service page shows what changed and when we could bracket it.
- Columns are alphabetical. There is no ranking, no score, no “best” — deliberately. What counts as acceptable coupling depends on your threat model, not ours. The launch findings walk through what the data shows; the methodology explains every marker and how evidence is verified.
Reuse#
The dataset (our metadata and annotations) is CC BY 4.0. Quoted provider text remains the provider’s, reproduced in short sanitized excerpts for verification. Machine-readable export: index.json.
Cite as: Cora Aegis, “KYC Drift” (v2026.07), cypherpunkguide.com/en/data/kyc-drift/, CC BY 4.0
Corrections and provider objections: editor@cypherpunkguide.com — errors are fixed and corrections recorded publicly.