Enterprise & Industry

Stanford's AI Observatory exposes blind spots in Anthropic, OpenAI usage data

Independent study finds 48% of AI conversations are work-unrelated, contradicting corporate reports.

Deep Dive

AI companies like Anthropic and OpenAI publish usage reports, but these only show what they choose to disclose, with no external verification. Stanford's AI Observatory, co-led by Anka Reuel of the Trustworthy AI Research Lab, aggregates and analyzes real user conversations from seven consent-based datasets spanning models such as Claude, Gemini, Grok, and ChatGPT. When researchers applied Anthropic's work-focused methodology to their data, 48% of conversations were filtered out. These excluded chats were far more likely to involve health and relationships (44.2% vs. 31.2%), adult or illicit topics (7.9% vs. 2.1%), harassment (27.5% vs. 5.66%), and sexual content (16.7% vs. 2.4%). The project aims to fill a critical gap for researchers and policymakers making consequential decisions about AI risks and benefits.

The Observatory also found significant differences across platforms. Grok was used heavily for news and politics but concentrated misinformation; Claude dominated coding, Gemini led in social and roleplay use, and ChatGPT was top for homework help. Over 2023–2025, conversations grew longer with more small talk, suggesting rising AI companionship, while sensitive exchanges declined—a sign of improving safeguards. Exposing these patterns in a unified, independent analysis, rather than separate corporate reports, gives a clearer picture of how AI is truly being used, according to researchers like David Widder of UT Austin.

Key Points
  • Stanford's AI Observatory analyzed 7 consented conversation datasets from Claude, Gemini, Grok, and ChatGPT
  • 48% of real conversations were filtered out by Anthropic's work-only methodology, hiding personal uses like health and relationships
  • Model behavior diverged sharply: Grok for news (with misinformation), Claude for coding, Gemini for roleplay, ChatGPT for homework

Why It Matters

Independent usage data helps regulators and researchers make evidence-based decisions on AI safety and policy.

📬 Get the top 10 AI stories daily