EmuGEMM: New fused kernels boost NVIDIA matrix math by up to 5.5x
EmuGEMM sustains 3,654 Tops on Blackwell, beating cuBLAS by 1.7x on TF32
Modern GPUs dedicate increasing silicon to low-precision matrix units, but scientific computing demands high precision. Existing Ozaki Scheme implementations waste performance by repeatedly writing intermediate results to global memory. EmuGEMM solves this with fused integer Tensor Core kernels that keep all intermediate data on-chip, eliminating the memory bottleneck. On NVIDIA Hopper GPUs, EmuGEMM sustains 1,639 Tops using Scheme I (83% of INT8 theoretical peak) and 3,654 Tops on Blackwell (81% of INT8 peak). This translates to 1.4x faster throughput than cuBLAS TF32 on Hopper and 1.7x on Blackwell, at equivalent accuracy.
For complex arithmetic, EmuGEMM's Scheme II delivers even larger gains: 2.3x over cuBLAS ZGEMM on Hopper and 5.5x on Blackwell. The paper, authored by Denghui Lu and colleagues at ETH Zurich (arXiv:2606.25453), demonstrates that fused kernels can dramatically narrow the precision-throughput gap for scientific workloads. By offloading emulated high-precision GEMM entirely to INT8 Tensor Cores, EmuGEMM opens the door to faster AI training with higher precision, dense linear algebra solvers, and physics simulations—all without sacrificing the accuracy that scientific users require.
- EmuGEMM achieves 1,639 Tops on Hopper (83% of INT8 peak) and 3,654 Tops on Blackwell (81% of INT8 peak) for Ozaki Scheme I
- Outperforms cuBLAS TF32 by 1.4x on Hopper and 1.7x on Blackwell with comparable accuracy
- For complex arithmetic (Scheme II), beats cuBLAS ZGEMM by up to 2.3x on Hopper and 5.5x on Blackwell
Why It Matters
Fused Tensor Core kernels close the precision-throughput gap, accelerating scientific computing and high-precision AI workloads on modern GPUs.