ContactWorld shows point-cloud + tactile boosts robot manipulation to 36.1%
New benchmark reveals best sensory combo for delicate robotic tasks like screwing and insertion.
ContactWorld, introduced by Zhiyuan Zhang and eight co-authors, is a systematic empirical study and benchmark designed to answer what makes vision-tactile world models effective for contact-rich manipulation. Across 12 tasks — including insertion, disassembly, screwing, and exploratory interaction — the team tested various sensory representations. Their key finding: spatially structured and temporally continuous representations consistently outperform others. Point-cloud observations raised average planning success from 20.7% (wrist-view) or 22.0% (front-view) to 32.1%. Adding tactile force-field representations, which preserve richer spatial structure and interaction dynamics, pushed performance further to 36.1%.
The research also highlights that tactile sensing's effectiveness depends crucially on cross-modal compatibility, not just on scaling the number of sensors. Under long-horizon planning objectives, where prediction errors and contact uncertainty accumulate, tactile input becomes increasingly important. These insights provide concrete guidance for roboticists building world models: prioritize representation structure and multimodal alignment over raw sensor quantity. The work is particularly relevant for industrial automation, assembly, medical robotics, and any domain requiring precise, sustained physical contact with objects.
- Point-cloud observations improved planning success from ~20% to 32.1% across 12 contact-rich tasks.
- Combining point-cloud with tactile force-field representations further boosted performance to 36.1%.
- Tactile sensing becomes significantly more valuable under long-horizon planning due to compounding prediction errors.
Why It Matters
Enables robots to perform precise assembly and maintenance tasks with fewer failures, advancing industrial automation and surgical robotics.