June 2026 AI benchmarks: Gemini 3.1, Claude Fable 5, GPT-5.5 vie for top spots
Gemini 3.1 Pro Preview leads hardest exam; Claude Fable 5 tops common sense.
Deep Dive
Key Points
- Gemini 3.1 Pro Preview achieves 46.4% on Humanity's Last Exam, the highest among 10+ top models.
- Claude Fable 5 leads SimpleBench with 81.9%, outperforming GPT-5.5 Pro by 5 percentage points.
- Independent benchmarks (Epoch, Scale) provide unbiased scores, unlike vendor self-reports.
Why It Matters
Independent benchmarks reveal which models truly handle complex reasoning and common sense, guiding enterprise AI selection and investment.