Developer Tools

NVIDIA Nemotron 3.5 Lightning accelerates agents 4x with NeMo Switchyard

NVIDIA's 30B MoE model slashes agent task times by 30%—plus a new open-source model router.

Deep Dive

NVIDIA is expanding its Nemotron 3 family with Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts (MoE) model purpose-built for high-volume, always-on agentic workloads. It claims up to 4x faster output speed and 30% faster agentic task completion than other models in its class, while maintaining frontier-level accuracy on PinchBench benchmarks. As an open model, it can be post-trained with NVIDIA NeMo on an organization's own domain data, tools, and workflows. The model also comes with Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset for coding-agent capabilities. Enterprises retain control over privacy and deployment, running Lightning locally on RTX PCs, DGX Spark, DGX Station, Jetson, or scaling across edge devices, RTX PRO workstations, data centers, and cloud environments.

Alongside Lightning, NVIDIA released NeMo Switchyard, an open-source library for smart routing inside popular agent tools. NeMo Switchyard lets enterprises build custom routers that direct each request to the most capable and suitable model—across a mix of open, proprietary, and NVIDIA models—without rewriting applications. This system-of-models approach reflects NVIDIA's vision: a frontier reasoning model like Nemotron 3 Ultra plans and orchestrates, while specialized models like Lightning handle code review, security alerts, billing, and tool use. Early adopters include CrowdStrike for cybersecurity, Harvey with Trajectory for legal services, CodeRabbit with Baseten for code review, Lila Sciences for life sciences reasoning, and Fastino Labs, which saw leading accuracy in software development, finance, and healthcare workloads.

Key Points
  • Nemotron 3.5 Lightning is a 30B-parameter MoE open model delivering up to 4x faster output and 30% faster agentic task completion than peers.
  • NeMo Switchyard is a new open-source router that dynamically sends requests to the best model across open, proprietary, and NVIDIA families without code rewrites.
  • Runs anywhere from RTX PCs and DGX systems to data centers and cloud, with customization partner results from CrowdStrike, Harvey, CodeRabbit, Lila Sciences, and Fastino Labs.

Why It Matters

Enterprises gain faster, private, and controllable agentic AI with efficient model routing—cutting latency and infrastructure costs.

📬 Get the top 10 AI stories daily