Open Source

Artificial Analysis accused of rigging AI index to favor Anthropic

Open-source Qwen loses top spot after index weights quietly changed

Deep Dive

Artificial Analysis (AA), a widely cited AI benchmarking platform, has come under fire after updating its intelligence index to v4.1.1. Reddit user Infinite-Local5435 publicly accused AA of deliberately tweaking the weights of the GDPval and T3 Banking evaluations so that Anthropic's Opus would once again rank above the open-source Qwen 3.8 Max. According to the post, Qwen had previously held an 8% lead on the T3 Banking metric, while Opus only managed a 5% advantage on GDPval—a mismatch that, critics argue, makes the re-ranking statistically suspicious.

The accusations go beyond simple version changes. The user claims AA "just so happen to launch" the update right after Qwen claimed the top spot on the agentic index, and speculates the move may have been financially motivated. "Highly likely to be paid off imo," the post reads. Others on the subreddit note that the pre-change scores were visible before v4.1.1, making the weight adjustment easy to verify. If proven, this would be a major blow to AA's credibility as an independent benchmark—a critical resource for developers and enterprises choosing between proprietary and open-source models.

Key Points
  • AA's v4.1.1 index update changed GDPval and T3 Banking weights, demoting Qwen 3.8 Max below Anthropic's Opus.
  • Qwen reportedly led T3 Banking by 8%, while Opus only led GDPval by 5%, raising statistical red flags.
  • Reddit users accuse Artificial Analysis of accepting payment to favor Anthropic, challenging the platform's neutrality.

Why It Matters

If benchmarks are biased, developers and enterprises can't trust AI comparisons, affecting model selection.

📬 Get the top 10 AI stories daily