SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC
SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC
Deep Dive
Computer Science > Distributed, Parallel, and Cluster Computing arXiv:2609.18110 (cs) [Submitted on 16 Sep 2026] Title: SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC Authors: Fangzhou Liang , Yibin Shen , Jianmin Hu , Jiayang Xu , Hanchi Gao , Minxian Xu ,