Nvidia's Rubin AI platform claims 10x efficiency over Grace Blackwell
Nvidia promises 10x tokens-per-watt with 72 GPUs per rack
Nvidia has unveiled concrete performance metrics for its next-generation Rubin AI platform, claiming it will deliver up to 10 times the tokens-per-watt efficiency compared to the existing Grace Blackwell system. The announcement, made on July 22, 2026, underscores Nvidia's strategic shift from a GPU supplier to a comprehensive AI computing systems provider. The centerpiece is the NVL72 rack, which pairs 72 Rubin GPUs with 36 Vera CPUs in a tightly integrated architecture designed for hyperscale AI training and inference.
This efficiency leap comes from architectural improvements in both the Rubin GPU and the Vera CPU, along with advanced interconnects that minimize data movement overhead. By achieving 10x tokens-per-watt, Nvidia aims to dramatically lower the total cost of ownership for large language model deployments, making AI infrastructure more accessible while reducing power consumption. The Rubin platform is expected to power next-generation data centers, competing directly with cloud custom silicon and other accelerators. Early benchmarks suggest that workloads currently requiring multiple Grace Blackwell racks could be condensed into a single Rubin rack, saving space, cooling, and energy.
- Rubin AI platform achieves 10x tokens-per-watt efficiency vs Grace Blackwell
- Each NVL72 rack integrates 72 Rubin GPUs and 36 Vera CPUs
- Nvidia positions itself as a full-stack AI computing systems provider
Why It Matters
10x efficiency means cheaper, greener AI training and inference at scale, reshaping data center economics.