Claude Opus 5 challenges GPT-5.6 Sol with 1M context and stronger reasoning
Claude Opus 5 matches GPT-5.6 on reasoning, sweeps agentic benchmarks at 1M context.
Anthropic's Claude Opus 5, the flagship model codenamed Honeycomb, officially launched on July 24, 2026, after a brief delay from the expected July 23 release. It features a 1M-token context window and configurable effort settings ranging from low to max, with a Fast Mode research preview that delivers 2.5x speed at 2x cost. Pricing is set at $5 per million input tokens and $25 per million output tokens, matching Opus 4.8 and undercutting Fable 5 by 50%. Early benchmark results show Opus 5 leading on several key metrics: it scores 30.2% on ARC-AGI-3 (novel problem solving) compared to GPT-5.6 Sol's 7.8%, and 70.6% on OSWorld 2.0 (computer use) versus Sol's 62.6%. On Frontier-Bench (terminal coding), Opus 5 achieves 43.3%, more than doubling Opus 4.8's 21.1%. However, GPT-5.6 Sol retains the lead on DeepSWE with 72.7%. The Intelligence Index rates Opus 5 (max) at 61, narrowly ahead of Fable 5 (60) and Sol (59), at 26% lower cost per task than Fable 5.
Anthropic positions Opus 5 as a daily driver for long-running agents, advanced coding, and professional knowledge work. It includes safety fallbacks that resolve to Opus 4.8 when needed. The model is available on Claude API, Claude Code, Pro/Max/Team/Enterprise, AWS Bedrock, and Google Vertex AI. Hands-on impressions are positive: Alex Albert notes greater token efficiency than Fable 5 for coding tasks, Ivan Fioravanti calls it "another level" versus Opus 4.8 or Sonnet 5, and Max Weinbach says Opus 5 and GPT-5.6 Sol feel "basically the same model" in daily useβthe benchmark gap being wider than the felt gap. GPT-5.6, meanwhile, continues as OpenAI's incumbent with three variants (Sol, Terra, Luna), Ultra Mode reasoning, and multi-agent APIs, plus Codex folded into ChatGPT. The competition between the two is now tighter than ever, with Opus 5 offering a compelling cost-performance advantage for reasoning-heavy workloads.
- Claude Opus 5 (codenamed Honeycomb) launched July 24, 2026 with a 1M-token context window and adjustable effort settings from low to max.
- Outperforms GPT-5.6 Sol on ARC-AGI-3 (30.2% vs 7.8%), OSWorld 2.0 (70.6% vs 62.6%), and Frontier-Bench (43.3% vs 21.1% for Opus 4.8).
- Fast Mode offers 2.5x speed at 2x cost; pricing at $5/$25 per M tokens matches Opus 4.8 and is half of Fable 5.
Why It Matters
Anthropic and OpenAI are neck-and-neck; Opus 5 offers strong reasoning at half the cost of Fable 5.