Developer Tools

AMD buys Taalas to make AI inference 48x faster

AMD acquires Taalas to etch AI models directly into silicon, hitting 17,000 tokens/sec.

Deep Dive

AMD has acquired Taalas, a Toronto-based AI chip startup, to supercharge inference performance by embedding model weights directly into silicon. The move signals AMD’s aggressive push into the premium AI hardware market, traditionally dominated by Nvidia. Taalas’ approach abandons conventional GPUs and dataflow architectures like Groq’s LPUs or Cerebras’ wafer-scale accelerators, instead using model-specific integrated circuits (MSICs). Their first test chip, the HC1, was fabricated on TSMC’s 6nm process and demonstrated blazing speed: 16,960 tokens per second on Meta’s Llama 3.1 8B, a 48x improvement over Nvidia’s GPUs and 8.5x faster than Cerebras’ accelerators at the time of its February unveiling.

The technology relies on two key components: a mask-ROM recall fabric for etched model weights and an SRAM recall fabric for KV caches and fine-tuning adapters. Taalas’ second-gen HC2 chip, expected this summer, aims to support up to 20 billion parameters per chip, enabling efficient scaling to trillion-parameter models via pipeline parallelism. For example, just 50 of these chips could support a 1-trillion-parameter model—far more space and power-efficient than Nvidia’s recent LPX systems. AMD plans to integrate Taalas’ accelerators with its Instinct-based Helios racks, creating a disaggregated architecture where GPUs handle compute-heavy prompt processing while token generation shifts to Taalas-based accelerators.

Key Points
  • AMD acquires Taalas to disrupt AI inference with silicon-etched model weights, achieving up to 17,000 tokens/sec (48x faster than Nvidia GPUs).
  • Taalas’ HC1 chip (6nm TSMC) ran Llama 3.1 8B at 16,960 tokens/sec, while HC2 targets 20B parameters per chip for scalable trillion-parameter models.
  • AMD will pair Taalas’ accelerators with Instinct-based Helios racks, enabling disaggregated AI workloads but locking users into specific models.

Why It Matters

AMD’s Taalas acquisition could redefine high-performance AI inference, making ultra-fast token generation cheaper and more efficient—but at the cost of model flexibility.

📬 Get the top 10 AI stories daily