Open Source

Cohere's unreleased 30B coding model lets devs test locally

Test Cohere's new code model on Hugging Face—only 3B active params for fast local inference.

Deep Dive

Cohere, the AI startup known for enterprise language models, is venturing into code generation for the first time. In a Reddit post, co-founder Nick Frosst offered the r/LocalLLaMA community early access to an unreleased coding model, calling it "small but fast." The model weighs 30 billion parameters in total but activates only 3 billion per token—likely through a Mixture-of-Experts architecture—making it suitable for running on consumer hardware. Frosst specifically invited developers to test the model on their own machines and report back, saying the team wants to "build from our learnings with this release." The weights are live on Cohere's Hugging Face page, though Frosst noted the model isn't fully ready for public launch and encouraged testers to focus on their own use cases.

Early benchmarks indicate token generation speeds are in line with other models of similar size and architecture, which is notable given the 10x active-parameter compression. The model is Cohere's first dedicated coding model, signaling the company's expansion beyond text generation into developer tools. By releasing it early to a knowledgeable community, Cohere is betting on rapid, targeted feedback to improve performance before a wide launch. The move mirrors strategies used by other AI labs, but the emphasis on local inference and transparent community engagement sets it apart. For developers, this is a rare chance to influence a production model's direction and test cutting-edge code generation without cloud dependencies.

Key Points
  • 30B parameters total, only 3B active per token for efficient local inference.
  • Available now on Hugging Face for early testing before official launch.
  • Cohere seeks community feedback to refine the model's code generation capabilities.

Why It Matters

Cohere's first coding model opens local AI development to devs, leveraging community feedback for a better production release.

📬 Get the top 10 AI stories daily