Research & Papers

How a New AI System Gives E-Commerce Video Feeds a 50% Boost — Without Any User History

How CLIP-based zero-shot retrieval solves the extreme cold-start problem for short-form video feeds...

Deep Dive

E-commerce platforms are shifting from static search to dynamic video feeds, but new short-form videos face an 'extreme cold-start' problem — no interaction history for collaborative filtering, plus strong position and duration biases. Researchers present VCG (Video Candidate Generation), a scalable multimodal retrieval engine that uses a domain-adapted CLIP model to map users and videos into a shared semantic space, enabling zero-shot retrieval based on visual content alone.

In rigorous evaluation, VCG compared generative (LLM) vs. discriminative (CLIP) embeddings. Generative models excelled at attribute prediction but suffered from embedding space collapse in retrieval tasks. VCG's CLIP-based approach mitigated engagement biases, yielding a 50% uplift in deep video completion in online A/B tests. The system also features three interactive retrieval scenarios: Product-to-Video, Video-to-Product, and Zero-Shot Semantic Search, demonstrating practical deployment for live e-commerce feeds.

Key Points
  • VCG uses domain-adapted CLIP for zero-shot video retrieval without requiring interaction history.
  • Generative LLM embeddings caused embedding space collapse, making them unsuitable for retrieval tasks.
  • Online A/B testing showed a 50% uplift in deep video completion rates with VCG.

Why It Matters

Enables personalized video feeds for new products without user history, a critical step for shoppable video commerce.

📬 Get the top 10 AI stories daily