Moonshot AI's Kimi K2.6 beats GPT-5.4 and Claude Opus, costs 5x less
Open-weight model with 256K context and 1T parameters quietly tops agent benchmarks
On April 20, Moonshot AI quietly released Kimi K2.6, an open-source model that now tops several key benchmarks. With a Mixture-of-Experts architecture (32B active parameters out of 1T total) and a 256K token context window, it outperforms GPT-5.4 and Claude Opus 4.6 on agent-centric tests like SWE-Bench Pro (58.6 vs 57.7) and Humanity's Last Exam with tools (54.0 vs 52.1/53.0). Its cost is dramatically lower — $0.55 per million input tokens vs. $3.00 for Claude Opus and $2.50 for GPT-4o — making it the most cost-effective frontier model available.
However, Kimi lags in pure mathematical reasoning (8-10 points behind Gemini 3.1 Pro) and Russian-language tasks. Independent verification of benchmarks is still limited, and the claimed 300-agent swarm capability remains unproven in production. Still, for Russian-speaking micro-businesses facing account bans and payment hurdles with Western APIs, Kimi offers a viable, open-weight alternative under a modified MIT license — a potential game-changer for affordable, unrestricted AI access.
- Kimi K2.6 scores 54.0 on Humanity's Last Exam (with tools) vs. GPT-5.4's 52.1 and Claude Opus 4.6's 53.0
- Costs $0.55 per million input tokens — 5-6x cheaper than GPT-4o ($2.50) and Claude Opus ($3.00)
- Features 256K token context window, 1T total parameters with 32B active per query, and MIT-like license for commercial use
Why It Matters
Cheaper, open-weight model rivals top Western AI, democratizing access for cost-sensitive enterprises and regions with payment restrictions.