Research & Papers

OneRec's OneReason adds chain-of-thought reasoning to generative recommendations

Generative rec models get reasoning capabilities inspired by LLMs' 'think before answer'

Deep Dive

Generative recommendation models from OneRec are widely used but have lacked reasoning ability because meaningful chain-of-thought (CoT) sequences cannot be constructed with only item tokens. Inspired by the ‘think before answer’ paradigm from LLMs, the team conducted preliminary studies (OneRec-Think, OpenOneRec) but found thinking mode showed no advantage over non-thinking. Drawing insights from CoT robustness in multimodal models, they identified two missing factors: perception (grounding item tokens in language semantics) and cognition (reorganizing user behavior sequences into latent interest points). This led to OneReason.

OneReason comprises three innovations: (1) strong itemic token perception during pre-training, (2) a three-level cognition-enhanced CoT format for supervised fine-tuning that structures reasoning into perception, cognition, and generation, and (3) a specialize-then-unify training recipe in reinforcement learning to strengthen thinking ability. The approach activates reasoning in models already deployed in real-world services like short-video, live-streaming, advertising, and e-commerce. The paper, authored by the OneRec Team (83 contributors), is a work in progress. If successful, OneReason could bring the interpretability and reasoning power of LLMs to recommendation systems, potentially improving personalization and accuracy at scale.

Key Points
  • OneReason enables chain-of-thought reasoning in generative recommendation, solving prior failures where thinking showed no advantage.
  • Success depends on both perception (grounding item tokens in language semantics) and cognition (reorganizing behavior into interest points).
  • Uses a three-stage pipeline: pre-training with strong perception, SFT with three-level CoT, and RL with specialize-then-unify training.

Why It Matters

Enables recommendation systems to reason like LLMs, potentially improving personalization and accuracy across major platforms.

📬 Get the top 10 AI stories daily