Skip to main content

KYC Drift: Hosted AI — Identity & Data-Handling Facts

Ask a simpler question than “is this AI anonymous?” — because that question has no honest yes/no answer. Ask instead: which identity anchors does each provider’s own paperwork attach to your use, and what does it say happens to what you type? This dataset records exactly that, for eight hosted-AI services, from their own terms, privacy policies, and help pages — each value linked to an archived copy of the document that states it.

The matrix
#

DimensionClaudeAnthropicDeepSeekDeepSeekGeminiGoogleLe ChatMistral AIChatGPTOpenAIOpenRouterOpenRouterPerplexityPerplexityVeniceVenice
Identity at signup & use What identity the provider requires/collects at each touchpoint.
Account requiredrequirednot assessedrequiredconditionalnot assessedrequirednot assessedoptional
Signup identifier(s)not assessedemail or phonerequirednot assessedone of: email / phone / SSOnot assessedrequirednot assessed
Phone at signuprequired email or phonenot assessedcollected if you provide itone of: email / phone / SSOnot assessednot assessednot assessed
Email at signupconditionalemail or phonenot assessedcollectedone of: email / phone / SSOconditionalnot assessednot assessed
Login methods (SSO/wallet)Google SSO or email magic-linknot assessednot assessednot assessednot assessednot assessednot assessedemail, social login, or Web3 wallet
Minimum age18+not assessednot assessed13+not assessednot assessednot assessednot assessed
Later ID/phone checkconditionalnot assessednot assessednot assessednot assessednot assessednot assessednot assessed
Payment & billing identitynot assessednot assessednot assessednot assessedrequiredacceptednot assessednot assessed
Data handling What happens to your prompts/data (stored, retained, trained on, reviewed, shared).
Prompts sent downstreamnot assessednot assessednot assessednot assessednot assessedsent to downstream providernot assessedzero data retention (imposed downstream)
Inputs used to trainon by default; opt-out availableon by default; no stated opt-outnot assessedon by default; opt-out availableon by default; opt-out availablevaries by downstream provider — selectableemail content excludednot disclosednot assessed
Human review of chatsnot assessednot assessedyesnot disclosednot assessednot assessednot assessednot assessed
Retentionnot assessednot assessedup to 3 yearsuntil you delete itnot assessednot assessednot assessednot retained
Deletion & identifier reuseavailableavailablenot assessedavailableavailablereusable after deletionnot assessednot assessednot assessed
Privacy lever (temp chat/ZDR/opt-out)not assessednot assessedavailableavailablenot assessednot assessednot assessednot assessed
Data-controller jurisdictionnot assessedChina (CN)not assessedFRnot assessednot assessednot assessednot assessed
Network layer (separate) OUT OF the provider-identity decoupling scope. Recorded as fact only. IP/device/analytics are network-layer, not provider-account identity.
Third-party site analyticsnot assessednot assessednot assessednot assessednot assessedpresentnot assessednot assessed
Metadata collected (IP etc.)not assessednot assessedcollectednot assessednot assessednot assessednot assessedcollected

Columns are alphabetical. · Every value is what the provider's own document states — not verified enforcement. · not assessed = not yet extracted from primary sources — an empty dimension is NOT evidence that no requirement exists · ⚠ = conflicting sources · ↻ = changed over time · No scores, no rankings, no recommendations.

How to read this
#

  • Every value is a provider statement. We record what the provider’s published documents say — we do not test enforcement, and a policy saying “we collect X” is not the same as a signup screen refusing to proceed without X. Where those differ, we record them separately.
  • “Not assessed” means exactly that. We have not yet extracted that dimension from primary sources. It is not evidence that no requirement exists — treating silence as “no requirement” is the single most common way comparison tables mislead, so this table never leaves a cell blank.
  • ⚠ marks a live conflict. When sources disagree (official help says one thing, a third-party report claims another), we show the official statement as the value and record the disagreement on the service page — we do not silently pick a side.
  • ↻ marks drift. Requirements change. Where we have evidence of an earlier state, the service page shows what changed and when we could bracket it.
  • Columns are alphabetical. There is no ranking, no score, no “best” — deliberately. What counts as acceptable coupling depends on your threat model, not ours. The launch findings walk through what the data shows; the methodology explains every marker and how evidence is verified.

Reuse
#

The dataset (our metadata and annotations) is CC BY 4.0. Quoted provider text remains the provider’s, reproduced in short sanitized excerpts for verification. Machine-readable export: index.json.

Cite as: Cora Aegis, “KYC Drift” (v2026.07), cypherpunkguide.com/en/data/kyc-drift/, CC BY 4.0

Corrections and provider objections: editor@cypherpunkguide.com — errors are fixed and corrections recorded publicly.