Research & Papers

Evolution Simulations Run Faster on 'Good Enough' Supercomputers

Supercomputers can now work through hardware failures, speeding up science for less money.

Deep Dive

Supercomputers are the heavy lifters of modern science, but they have a problem: with millions of parts, something is always breaking. Normally, a single glitch forces the whole machine to slow down or restart. Researchers asked: what if we just let those small errors happen and keep going?

They tested this on 'digital evolution' — software that mimics how nature evolves creatures, letting millions of virtual organisms adapt and solve hard problems. Using a relaxed approach called best-effort computing, they ran these simulations across 64 processors in a cluster. Even when hardware acted up, the system stayed reliable and delivered a 2.1-times speedup over the traditional method, with 92% scaling efficiency (meaning adding machines almost proportionally speeds things up).

The team also explored an even more extreme case: a massive chip called the Cerebras Wafer-Scale Engine, which packs 880,000 processors onto one silicon wafer. That many parts makes failures common. They showed that by sampling data occasionally and tolerating errors, simulations can still track the history of evolving populations — as long as you verify the results instead of blindly trusting them.

The big idea, say the authors, is that the future of supercomputing may be 'post-deterministic': machines that don't guarantee perfect, repeatable results but instead produce reliably good outcomes. Digital evolution, which is already used to design drugs, optimize robots, and solve complex engineering puzzles, is the perfect proving ground. The payoff could be cheaper, faster experimental science — and new discoveries that would have been too expensive or fragile to run before.

Key Points
  • Running evolution simulations on 'good enough' computers gave a 2.1x speedup across 64 processors, even with faulty hardware.
  • The approach works at 92% efficiency, meaning adding more machines still makes the work faster without waste.
  • The same tolerant strategy was applied to an 880,000-processor super-chip, showing it can handle enormous, error-prone systems.

Why It Matters

This could slash the cost of supercomputer research, making breakthroughs in medicine and engineering faster and more affordable.

📬 Get the top 10 AI stories daily