Moonshot AI's Kimi K3: 2.8T-parameter open-weight model claims world's largest
2.8 trillion parameters, 1M context, self-hostable at ~1.4TB – and cheaper than US rivals.
Moonshot AI dropped Kimi K3 this week, a 2.8-trillion-parameter sparse Mixture-of-Experts model that immediately becomes the world's largest open-weight AI system. The model packs 896 experts (about 16 active per forward pass), supports native multimodal input (text, image, video), and offers a 1-million-token context window. Two novel architectural innovations—Kimi Delta Attention (KDA) and Attention Residuals—handle long-context efficiency, while MXFP4 weight quantization shrinks storage to roughly 1.4 TB, making multi-node GPU self-hosting feasible for well-resourced teams. Pricing is notably aggressive: $3 per million input tokens and $15 per million output tokens, undercutting comparable U.S. frontier models. On benchmarks, Kimi K3 took first place in six of seven domains in the Frontend Code Arena and scored 88.3 on Terminal-Bench 2.1. Open weights will be released on July 27, 2026.
In parallel, Thinking Machines Lab shipped Inkling, a 975B-parameter open-weights multimodal MoE model under Apache 2.0, giving enterprises a Western alternative for data sovereignty with 1M context and native text/image/audio support. Meanwhile, France's competition authority issued a 3,700-page advisory warning that OpenAI, Google, and Anthropic control over 84% of the AI agent market—having built and deployed their own agents to audit the space. These developments underscore a pivotal week: open-weight frontier models are becoming practical self-hosted options, while regulators step up scrutiny on market concentration and privacy practices (the Grok Build CLI scandal, quietly uploading 5.1GB of dev data, adds urgency).
- Kimi K3: 2.8T total parameters, 896 experts (16 active), 1M token context, native vision.
- MXFP4 quantization reduces storage to ~1.4TB for practical self-hosting on multi-node GPU clusters.
- Charges $3/M input tokens ($15/M output) – significantly cheaper than comparable US frontier models.
- Achieved top scores in Frontend Code Arena (6/7 domains) and Terminal-Bench 2.1 (88.3).
Why It Matters
Self-hosting a frontier-quality model is now realistic, undercutting API costs and preserving data control for enterprises.