MiniMaxAI drops MiniMax-M3: 428B params, only 23B active
A massive 428B-parameter model that activates just 23B per token...
Deep Dive
Minimax m3 weights are out! It has ~428B parameters and ~23B activated parameters.
Key Points
- ~428B total parameters, only ~23B activated per inference via MoE architecture
- Weights openly released on Hugging Face for download and fine-tuning
- Efficiency gains allow serving a 400B+ scale model with far lower compute than dense alternatives
Why It Matters
MiniMax-M3 offers frontier-scale AI capability with a fraction of the compute cost, democratizing access.