Research & Papers

Neuromorphic MDLMs combine spike sparsity & block diffusion for faster inference

New N-MDLM model boosts token throughput while slashing energy with spike-based sparsity.

Deep Dive

Autoregressive large language models (LLMs) suffer from low operational intensity and high energy consumption because each generated token requires accessing the full set of model parameters. Masked diffusion language models (MDLMs) partially address this by allowing multiple tokens per parameter access, but still struggle on compute-bound platforms.

Now, a team from King's College London and other institutions introduces Neuromorphic MDLMs (N-MDLMs), which combine block diffusion with spike-based neuromorphic computation. The block diffusion technique generates several tokens in parallel from a single parameter fetch, while the spike-based sparsity mechanism skips inactive neural channels, drastically cutting effective parameter traffic and computations. In translation experiments, N-MDLMs showed significant gains in both energy efficiency and throughput—even outperforming standard MDLMs on compute-bound hardware. The work was accepted at the 2026 IEEE Workshop on Signal Processing Systems (SiPS) and points toward a future where sparsity-driven neuromorphic architectures could enable efficient on-device LLM inference.

Key Points
  • Block diffusion generates multiple tokens per parameter access, improving throughput.
  • Spike-induced sparsity skips inactive channels, reducing computational load and memory traffic.
  • Achieves energy efficiency gains even on compute-bound platforms where standard MDLMs fail to improve.

Why It Matters

Spike-based sparsity could unlock efficient large model inference on edge devices, reducing energy costs.

📬 Get the top 10 AI stories daily