Research & Papers

New Trick Trains AI Twice as Fast Using Quarter-Size Numbers

⚡Cheaper AI training could mean cheaper, faster chatbots for everyone.

Deep Dive

A new arXiv paper by Robert Hu presents "format-aware fusion," which co-designs each quantization producer with its scale domain and consumer layout for native mxfp4, global nvfp4, and cooperative-thread-array-local nvfp4, aiming to stop scale computation, operand packing, layout construction, and saved backward state from erasing the gains of FP4 Tensor Cores. In matched same-accelerator probes, bfloat16 and Transformer Engine nvfp4 reach 18.8K and 27.6K tokens/s/GPU, while the fastest custom route reaches 37.9K. Mxfp4 with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning reaches 37.2K tokens/s/GPU, 86.3% bfloat16 model FLOP utilization, and ends 2.11% above the raw bfloat16 training-loss endpoint. Downstream rankings differ from training-loss rankings, showing FP4 outcomes depend jointly on scale contract, operand, and execution path.

Key Points
  • Storing AI numbers in a tiny FP4 format (four bits instead of sixteen) roughly doubled training speed in a test on an 8-billion-parameter model.
  • In the fastest test, the model processed 37.9K words per second per chip, versus 18.8K with the standard 16-bit method — about 2x faster.
  • The speed came with a small quality penalty (about 2% worse training loss), and the paper is an unreviewed preprint, so real-world gains aren't guaranteed.

Why It Matters

Faster, cheaper AI training could lower the cost of the AI tools you already use every day.

📬 Get the top 10 AI stories daily