Image & Video

SoM-MTM: Masked token model repairs lost visual data in cooperative perception

A plug-and-play model recovers distorted features from packet loss using masked learning, improving perception.

Deep Dive

Cooperative perception (CP) in next-generation mobile networks lets multiple vehicles or agents share sensory data, but transmission over unreliable channels can drop packets and corrupt features. To solve this, researchers Haozhen Li, Rongqing Zhang, and Xiang Cheng — affiliated with Peking University's state key laboratory — introduce SoM-MTM (Synesthesia of Machines-driven Masked Token Model). The name echoes the way machines "feel" information across senses, and the model treats perception as a masked-image-modeling problem, similar to MAE. By learning to reconstruct missing tokens from context, SoM-MTM recovers distorted environmental features at the receiver end, effectively turning packet loss into a feature rather than a failure.

Technically, SoM-MTM is built on the Swin Transformer backbone and embeds prior masked information through an External Routing Mixture-of-Experts (MoE) mechanism, which dynamically selects expert modules to repair and enhance perception features during cooperation. The design is deliberately plug-and-play, meaning it can be dropped into existing visual cooperative perception pipelines without retraining from scratch. Experimental results demonstrate consistent improvement across various perception tasks — including those with unseen channel conditions and cooperation modes — while maintaining competitive model size and computational cost. This suggests that heavier, task-specific communication modules may no longer be necessary; a single, generic masked-token model can handle heterogeneous conditions. For autonomous driving, connected V2X systems, and edge-based robotics, SoM-MTM points to a future where AI agents share percepts reliably even over imperfect wireless links.

Key Points
  • SoM-MTM is a plug-and-play masked token model for cooperative perception, inspired by MAE to handle packet loss channels.
  • It uses a Swin Transformer backbone with an External Routing MoE to dynamically repair and enhance visual features.
  • Tests show consistent performance gains across diverse tasks and strong generalization to unseen scenarios with low model cost.

Why It Matters

Enables reliable AI cooperation over lossy networks—critical for autonomous driving and V2X, without bespoke communication models.

📬 Get the top 10 AI stories daily