Stanford's AI Observatory exposes blind spots in Anthropic, OpenAI usage data
Independent study finds 48% of AI conversations are work-unrelated, contradicting corporate reports.
AI companies like Anthropic and OpenAI publish usage reports, but these only show what they choose to disclose, with no external verification. Stanford's AI Observatory, co-led by Anka Reuel of the Trustworthy AI Research Lab, aggregates and analyzes real user conversations from seven consent-based datasets spanning models such as Claude, Gemini, Grok, and ChatGPT. When researchers applied Anthropic's work-focused methodology to their data, 48% of conversations were filtered out. These excluded chats were far more likely to involve health and relationships (44.2% vs. 31.2%), adult or illicit topics (7.9% vs. 2.1%), harassment (27.5% vs. 5.66%), and sexual content (16.7% vs. 2.4%). The project aims to fill a critical gap for researchers and policymakers making consequential decisions about AI risks and benefits.
The Observatory also found significant differences across platforms. Grok was used heavily for news and politics but concentrated misinformation; Claude dominated coding, Gemini led in social and roleplay use, and ChatGPT was top for homework help. Over 2023–2025, conversations grew longer with more small talk, suggesting rising AI companionship, while sensitive exchanges declined—a sign of improving safeguards. Exposing these patterns in a unified, independent analysis, rather than separate corporate reports, gives a clearer picture of how AI is truly being used, according to researchers like David Widder of UT Austin.
- Stanford's AI Observatory analyzed 7 consented conversation datasets from Claude, Gemini, Grok, and ChatGPT
- 48% of real conversations were filtered out by Anthropic's work-only methodology, hiding personal uses like health and relationships
- Model behavior diverged sharply: Grok for news (with misinformation), Claude for coding, Gemini for roleplay, ChatGPT for homework
Why It Matters
Independent usage data helps regulators and researchers make evidence-based decisions on AI safety and policy.