New Trick Trains AI Twice as Fast Using Quarter-Size Numbers
Cheaper AI training could mean cheaper, faster chatbots for everyone.
A new arXiv paper by Robert Hu presents "format-aware fusion," which co-designs each quantization producer with its scale domain and consumer layout for native mxfp4, global nvfp4, and cooperative-thread-array-local nvfp4, aiming to stop scale computation, operand packing, layout construction, and saved backward state from erasing the gains of FP4 Tensor Cores. In matched same-accelerator probes, bfloat16 and Transformer Engine nvfp4 reach 18.8K and 27.6K tokens/s/GPU, while the fastest custom route reaches 37.9K. Mxfp4 with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning reaches 37.2K tokens/s/GPU, 86.3% bfloat16 model FLOP utilization, and ends 2.11% above the raw bfloat16 training-loss endpoint. Downstream rankings differ from training-loss rankings, showing FP4 outcomes depend jointly on scale contract, operand, and execution path.
- Storing AI numbers in a tiny FP4 format (four bits instead of sixteen) roughly doubled training speed in a test on an 8-billion-parameter model.
- In the fastest test, the model processed 37.9K words per second per chip, versus 18.8K with the standard 16-bit method — about 2x faster.
- The speed came with a small quality penalty (about 2% worse training loss), and the paper is an unreviewed preprint, so real-world gains aren't guaranteed.
Why It Matters
Faster, cheaper AI training could lower the cost of the AI tools you already use every day.