Study: Grok's X deployment silently boosts pseudo-science scores 2-5x
Grok's Fast versions on X scored pseudo-science at 70-75 while others stayed at 15-40.
A new study from researchers Scarso, Almeida, and Pina (arXiv:2607.22513) reveals that how commercial LLMs are deployed—not just their underlying weights—dramatically affects their validation of pseudo-scientific claims. Testing four major LLM families (Claude, Grok, GPT, Gemini) on ethnonationalist pseudo-science derived from Frank Salter's biosocial framework, they found that Grok's Fast versions (which power the default X experience) consistently assigned credibility scores of 70-75—two to five times higher than all other models (which scored 15-40). This pattern was absent on control prompts about basic evolutionary consensus, where all models performed comparably.
Three additional findings compound the concern: (1) a silent patch reversed Grok's behavior from chaotic to stably high validation overnight, with zero public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and web (5.5) three months later; (3) the most defensible response—refusing to rate the pseudo-scientific claim—appeared sporadically in Claude Opus 4.1 (web) and GPT-5.1 Chat (API), then eroded in successor versions. The authors argue this 'opaque epistemic mediation' makes LLM behavior a matter of public concern, calling for new forms of accountability.
- Grok's Fast versions on X assigned pseudo-science credibility scores 70-75, 2-5x higher than Claude, GPT, and Gemini (15-40).
- A silent, undocumented patch reversed Grok's behavior from chaotic to stable high validation overnight.
- Same Grok model ID scored 75 via API but 5.5 via web three months later, showing interface routing matters.
Why It Matters
LLM behavior is not fixed; deployment configurations secretly shape scientific claims, demanding transparency from providers.