AI Safety

Study: Grok's X deployment silently boosts pseudo-science scores 2-5x

Grok's Fast versions on X scored pseudo-science at 70-75 while others stayed at 15-40.

Deep Dive

A new study from researchers Scarso, Almeida, and Pina (arXiv:2607.22513) reveals that how commercial LLMs are deployed—not just their underlying weights—dramatically affects their validation of pseudo-scientific claims. Testing four major LLM families (Claude, Grok, GPT, Gemini) on ethnonationalist pseudo-science derived from Frank Salter's biosocial framework, they found that Grok's Fast versions (which power the default X experience) consistently assigned credibility scores of 70-75—two to five times higher than all other models (which scored 15-40). This pattern was absent on control prompts about basic evolutionary consensus, where all models performed comparably.

Three additional findings compound the concern: (1) a silent patch reversed Grok's behavior from chaotic to stably high validation overnight, with zero public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) the most defensible response—refusing to rate the pseudo-scientific claim—appeared sporadically in Claude Opus 4.1 (web) and GPT-5.1 Chat (API), then eroded in successor versions. The authors argue this 'opaque epistemic mediation' makes LLM behavior a matter of public concern, calling for new forms of accountability.

Key Points
  • Grok's Fast versions on X assigned pseudo-science credibility scores 70-75, 2-5x higher than Claude, GPT, and Gemini (15-40).
  • A silent, undocumented patch reversed Grok's behavior from chaotic to stable high validation overnight.
  • Same Grok model ID scored 75 via API but 5.5 via web three months later, showing interface routing matters.

Why It Matters

LLM behavior is not fixed; deployment configurations secretly shape scientific claims, demanding transparency from providers.

📬 Get the top 10 AI stories daily