Robotics

ST-WAM boosts robot vision robustness by 21% in new model

New ST-WAM model from Tsinghua researchers achieves 92.8% accuracy on RoboTwin 2.0 and doubles real-world success under visual shifts

Deep Dive

A team of 15 researchers from Tsinghua University and other institutions introduced ST-WAM (Semantic-Temporal World Action Model), a breakthrough in robot manipulation that addresses a critical weakness in existing World Action Models (WAMs). These models traditionally struggle with "Training-Distribution Hallucination"—where robots imagine irrelevant training data instead of reacting to real-time visual changes. ST-WAM solves this by using DINOv3 as a shared semantic backbone for both future prediction and history retrieval, while preserving fine-grained VAE dynamics for action execution.

The model introduces two key innovations: Dual-Space Future Experts (DSFE) that predict both VAE latents and DINO features, and Current-Anchored Intent Retrieval (CAIR) that pulls task-relevant evidence from recent DINO history. Unlike prior approaches, ST-WAM requires no additional pretraining or task annotations and performs no explicit future generation during inference. Its effectiveness is proven through exceptional benchmark performance—98.7% on LIBERO and 92.8% on RoboTwin 2.0—while achieving a remarkable 21.3 percentage point improvement in zero-shot LIBERO-Plus performance and doubling real-world success rates under visual distribution shifts from 25.8% to 61.5%.

Key Points
  • ST-WAM from Tsinghua researchers achieves 98.7% accuracy on LIBERO and 92.8% on RoboTwin 2.0 benchmarks
  • Boosts zero-shot performance by 21.3 percentage points and doubles real-world success under visual shifts (25.8% → 61.5%)
  • Introduces Dual-Space Future Experts and Current-Anchored Intent Retrieval without requiring additional pretraining or annotations

Why It Matters

Solves critical robot hallucination problem, enabling reliable real-world manipulation in changing visual conditions

📬 Get the top 10 AI stories daily