NVIDIA's Nemotron 3 Ultra: Open MoE Hybrid Mamba-Transformer for Agentic Reasoning
Open-source model achieves 2x speed and 40% better reasoning on agentic tasks.
Deep Dive
A paper titled "Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning" authored by NVIDIA contributors was submitted to arXiv on June 12, 2026.
Key Points
- 200B total parameters with 40B active per token using MoE, achieving 2x faster inference than dense models
- Hybrid Mamba-Transformer architecture optimized for agentic reasoning (tool use, planning, code execution)
- Fully open-source release including weights, training code, data recipes, and custom agentic benchmarks
Why It Matters
Open-source hybrid model from NVIDIA democratizes efficient agentic AI, challenging proprietary models on reasoning tasks.