Research & Papers

Axolotl3D: Unified framework for faithful 3D shape completion from occluded images

Axolotl3D fuses point clouds, images, and visibility masks for robust 3D generation.

Deep Dive

A new paper from Anita Hu and Maria Shugrina, accepted to ECCV 2026, introduces Axolotl3D, a unified framework for faithful 3D shape completion that addresses a critical gap in existing generative models. While current 3D diffusion models excel at generating geometry from a single image, they assume complete visibility and single-view inputs, limiting their use in multi-view, occluded, or editing scenarios. Axolotl3D overcomes this by jointly conditioning on multiple inputs: images, visibility masks, camera parameters, and a partial point cloud. The point cloud serves as a geometric anchor, ensuring the completed shape remains faithful to the observed structure, while camera parameters align multiple views into a shared 3D coordinate system for consistent reconstruction.

The framework employs a unified training strategy that synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning even when some inputs are missing or occluded. On benchmark datasets Toys4K and OmniObject3D, Axolotl3D achieves state-of-the-art performance under both clean and occluded settings. It also shows strong results in real-world reconstruction and geometry-consistent editing, making it practical for applications like robotics, AR/VR, and content creation. By unifying multi-modal conditioning in a single model, Axolotl3D represents a step toward more controllable and robust 3D generation from imperfect real-world data.

Key Points
  • Axolotl3D jointly conditions on images, visibility masks, camera parameters, and partial point clouds for occlusion-aware 3D completion.
  • Achieves state-of-the-art performance on Toys4K and OmniObject3D benchmarks under both clean and occluded settings.
  • Enables geometry-consistent editing and real-world reconstruction, accepted to ECCV 2026.

Why It Matters

Axolotl3D bridges the gap between single-view generation and real-world occluded scenes, enabling practical 3D reconstruction for robotics and AR.

📬 Get the top 10 AI stories daily