Research & Papers

ByteDance's TM20K boosts ad revenue by 1.036% with 20K token sequences

ByteDance's TM20K model handles 20K-token sequences for ad recommendations with minimal latency cost...

Deep Dive

ByteDance's research team has unveiled TM20K, a groundbreaking framework for ultra-long sequence modeling in e-commerce ad recommendation systems. Published on arXiv (ID:2608.07055), TM20K introduces a two-stage knowledge distillation approach where a 'teacher' model retains full token attention while 'student' models use efficient token merging techniques. This architecture enables sequence lengths of 20,000 tokens—far exceeding typical recommendation systems—while maintaining near-identical training and serving costs to current state-of-the-art models.

The deployment in ByteDance's live e-commerce advertising system has already shown measurable business impact. Despite the massive increase in sequence length, the system delivered a 1.036% improvement in Ad Space Success (ADSS) metrics while only increasing serving latency by 5.6%. The framework replaces traditional sequence compression methods that lose fine-grained user interest information, instead preserving full transformer modeling capabilities through knowledge distillation. This represents one of the first successful large-scale implementations of ultra-long sequence modeling in production advertising systems.

Key Points
  • TM20K extends e-commerce sequence modeling to 20,000 tokens while keeping training/serving costs nearly identical to prior models
  • Deployed in ByteDance's live ad recommendation system, delivering a 1.036% ADSS improvement with only 5.6% latency increase
  • Uses two-stage knowledge distillation: full-attention teacher model + token-merged student models for efficiency

Why It Matters

Proves ultra-long sequence modeling can drive real revenue gains in production systems without proportional cost increases

📬 Get the top 10 AI stories daily