Research & Papers

AMD's STEEL brings FlashAttention to NPUs with 9x energy savings

New open-source engine slashes energy 9x for long-sequence LLM inference on laptops.

Deep Dive

Researchers from the paper's author list introduce STEEL, the first open-source FlashAttention implementation for AMD's XDNA NPUs. On the Ryzen AI 9 HX 370 SoC, STEEL reduces energy consumption by an average of 9.17x vs CPU and 1.75x vs GPU. It achieves a 9.6x latency reduction on XDNA 1 and a 22.8x speedup on XDNA 2 over a layer-by-layer attention implementation. Its sparsity-aware pipeline placement efficiently handles causal mask load imbalance.

Key Points
  • 9.17x energy reduction vs CPU and 1.75x vs GPU on AMD Ryzen AI 9 HX 370
  • 22.8x speedup on XDNA 2 NPU vs naive layer-by-layer attention
  • First open-source FlashAttention implementation targeting XDNA-like NPUs

Why It Matters

Enables privacy-preserving, energy-efficient LLM agents on AMD laptops without cloud dependence.

📬 Get the top 10 AI stories daily