Research & Papers

Nvidia's New Chips Get an 8% Speed Boost From Simple Math

⚡Faster, cheaper AI training could mean lower subscription prices — and quicker new models.

Deep Dive

Every time an AI model runs, it does a handful of specific math operations over and over. Two of the most common are sigmoid and tanh — think of them as S-shaped curves that squash any number into a small, tidy range. Chips have dedicated hardware for these, but that hardware is slow compared to ordinary multiplication. A researcher named Robert Hu tried replacing them with short polynomial programs, which are basically simple multiply-and-add recipes a calculator could follow.

He tested this on Nvidia's newest GB200 hardware, the chips used to train today's biggest AI models. In isolated tests, the simple recipes ran 1.19 to 2.19 times faster when the data sat in fast on-chip memory. In realistic training jobs, swapping in the shortcuts sped up a full training step by 2.7%, 2.9%, and 8.0% depending on the task, with one attention task gaining 7.4% on its forward pass.

Why should you care? Training a frontier AI model costs tens of millions of dollars in electricity and chip time. An 8% speedup is like getting a free extra workday every two weeks — it lowers the cost of every model built afterward, which eventually shows up as cheaper AI subscriptions and faster releases. Just as important, the models didn't get dumber: after about 100 billion words of training, the difference in learning quality was tiny, ranging from a hair better to a hair worse.

The catch is that this is a single-author preprint — not yet checked by other scientists. The gains vary a lot by task (most were under 3%, not 8%), they're tuned to one specific family of Nvidia chips, and quality was only measured over relatively short training runs. It's a useful shortcut, not a breakthrough that changes what AI can do.

Key Points
  • A researcher replaced three slow math functions inside AI models with short, simple arithmetic recipes that computers handle much faster.
  • Full AI training steps got 2.7% to 8% faster on Nvidia's newest GB200 chips — meaningful savings when training costs millions.
  • Model quality barely changed after 100 billion words of training, but the work is a single unverified paper tied to specific hardware.

Why It Matters

Cheaper, faster AI training today usually becomes lower prices and quicker new features for you tomorrow.

📬 Get the top 10 AI stories daily