Research & Papers

OneRec's OneReason adds chain-of-thought reasoning to generative recommendations

⚡Generative rec models get reasoning capabilities inspired by LLMs' 'think before answer'

Deep Dive

Generative recommendation models from OneRec are widely used but have lacked reasoning ability because meaningful chain-of-thought (CoT) sequences cannot be constructed with only item tokens. Inspired by the ‘think before answer’ paradigm from LLMs, the team conducted preliminary studies (OneRec-Think, OpenOneRec) but found thinking mode showed no advantage over non-thinking. Drawing insights from CoT robustness in multimodal models, they identified two missing factors: perception (grounding item tokens in language semantics) and cognition (reorganizing user behavior sequences into latent interest points). This led to OneReason.

OneReason comprises three innovations: (1) strong itemic token perception during pre-training, (2) a three-level cognition-enhanced CoT format for supervised fine-tuning that structures reasoning into perception, cognition, and generation, and (3) a specialize-then-unify training recipe in reinforcement learning to strengthen thinking ability. The approach activates reasoning in models already deployed in real-world services like short-video, live-streaming, advertising, and e-commerce. The paper, authored by the OneRec Team (83 contributors), is a work in progress. If successful, OneReason could bring the interpretability and reasoning power of LLMs to recommendation systems, potentially improving personalization and accuracy at scale.

Key Points
  • OneReason enables chain-of-thought reasoning in generative recommendation, solving prior failures where thinking showed no advantage.
  • Success depends on both perception (grounding item tokens in language semantics) and cognition (reorganizing behavior into interest points).
  • Uses a three-stage pipeline: pre-training with strong perception, SFT with three-level CoT, and RL with specialize-then-unify training.

Why It Matters

Enables recommendation systems to reason like LLMs, potentially improving personalization and accuracy across major platforms.

📬 Get the top 10 AI stories daily