Models & Releases

June 2026 AI benchmarks: Gemini 3.1, Claude Fable 5, GPT-5.5 vie for top spots

Gemini 3.1 Pro Preview leads hardest exam; Claude Fable 5 tops common sense.

Deep Dive

Key Points
  • Gemini 3.1 Pro Preview achieves 46.4% on Humanity's Last Exam, the highest among 10+ top models.
  • Claude Fable 5 leads SimpleBench with 81.9%, outperforming GPT-5.5 Pro by 5 percentage points.
  • Independent benchmarks (Epoch, Scale) provide unbiased scores, unlike vendor self-reports.

Why It Matters

Independent benchmarks reveal which models truly handle complex reasoning and common sense, guiding enterprise AI selection and investment.

📬 Get the top 10 AI stories daily