Google Gemini 2.5 Pro tops AI leaderboards by 40 ELO points
Google's new reasoning model beats all rivals with massive margin on LMArena
Google DeepMind dropped a bombshell this week with Gemini 2.5 Pro Experimental, a thinking model that now dominates the LMArena leaderboard by a staggering 39-40 ELO points—the largest gap ever seen. The model achieves #1 across multiple SEAL benchmarks including Humanity's Last Exam, VISTA (multimodal), Tool Use, and MultiChallenge. Available immediately in Google AI Studio and the Gemini App at no cost, it features unified reasoning, long context, and tool use. Google also began rolling out real-time AI video features (Astra) to Gemini Live, enabling conversational screen sharing and camera-based interactions—a major step toward ambient computing.
Beyond Gemini, the AI agent ecosystem heated up. Browser Use, an open-source tool that helps AI agents navigate websites using plain-text instructions, raised $17M in seed funding. The concept echoes Ethan Mollick's observation that "clearly written instruction manuals are the future of APIs." Separately, OpenAI CEO Sam Altman confirmed MCP (Model Context Protocol) support is coming to ChatGPT desktop and the Responses API, while Microsoft announced MCP integration for Copilot Studio—letting agents automatically stay updated without manual maintenance. Chinese startup Manus AI is reportedly seeking a $500M valuation, and PwC launched its own "AI Agent Operating System" for enterprises. The week made clear that 2025 is the year agents and reasoning models go mainstream.
- Google's Gemini 2.5 Pro Experimental is #1 on LMArena by a 39-40 ELO margin, beating all competitors
- Tool makes AI web navigation easier with plain-text instructions; raised $17M in seed funding
- OpenAI and Microsoft both announced support for MCP, enabling persistent, auto-updating agent connections
Why It Matters
Google reclaims AI leadership with a reasoning model that's free today, while agent infrastructure rapidly matures.