Research & Papers

ICML 2026: DIRECT Framework Lets You Insert 3D Objects with Pose Control

Researchers achieve pose-controllable object insertion with decomposed visual proxies.

Deep Dive

Object insertion — seamlessly placing a reference object into a background image — has long been a challenge in computer vision. Existing diffusion-based methods treat it as a 2D inpainting task, offering no control over the object's 3D orientation. Now, a team of researchers from Nanyang Technological University and other institutions presents DIRECT (Decomposed Injection for Reference Composition and Target-integration), a framework that integrates interactive pose manipulation with high-fidelity 2D synthesis. DIRECT breaks down the insertion conditions into three complementary components: appearance guidance (visual details from the reference), geometry guidance (derived from a user-adjusted 3D proxy), and context guidance (from the target background). These are injected through separate pathways, preventing feature entanglement and allowing the model to simultaneously preserve the object's appearance, follow the specified pose, and blend into the scene. The framework also includes an automated data construction pipeline to improve training data diversity and quality.

Accepted at ICML 2026, DIRECT outperforms previous methods on both geometric controllability and visual quality. For example, it handles complex poses and lighting changes that earlier approaches fail to manage. The system's intuitive interface lets users rotate, scale, or reposition objects in 3D space before final rendering. This work has direct implications for image editing, AR/VR content creation, and e-commerce — where accurate product placement with proper perspective is critical. By decoupling pose and appearance, DIRECT moves beyond simple 2D pasting toward truly 3D-aware composition, setting a new standard for controllable object insertion.

Key Points
  • Framework decomposes insertion into appearance, geometry, and context guidance channels to avoid feature entanglement
  • Users can interactively adjust the object's 3D pose (rotation, scale, position) before final compositing
  • Outperforms prior diffusion-based methods in both geometric controllability and visual quality benchmarks

Why It Matters

Enables realistic, pose-aware object insertion for professional image editing, AR/VR, and e-commerce product visualization.

📬 Get the top 10 AI stories daily