Research & Papers

SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC

SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC

Deep Dive

Computer Science > Distributed, Parallel, and Cluster Computing arXiv:2609.18110 (cs) [Submitted on 16 Sep 2026] Title: SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC Authors: Fangzhou Liang , Yibin Shen , Jianmin Hu , Jiayang Xu , Hanchi Gao , Minxian Xu ,

📬 Get the top 10 AI stories daily