Anthropic's Fable 5 underperforms below Gemini 3.1 on LiveBench — benchmark integrity questioned
New model from Anthropic scores lower than Google's Gemini 3.1, sparking debate on benchmark reliability.
Deep Dive
Is this benchmark broken, or is Anthropic benchmaxing? A Reddit post by user MohMayaTyagi asks whether a test called LiveBench is flawed or if Anthropic is optimizing for it.
Key Points
- Anthropic's Fable 5 scored below Google's Gemini 3.1 on the LiveBench benchmark, contradicting prior performance claims.
- Community members question whether LiveBench is broken or if Anthropic engaged in benchmaxing (over-optimizing for specific tests).
- Exact scores are undisclosed, but reasoning tasks reportedly caused the biggest performance gap.
Why It Matters
Questions over benchmark reliability force professionals to re-evaluate model rankings and demand more rigorous evaluation standards.