Research & Papers

GIRAF diffusion model generates realistic full-body interactions with articulated objects

New AI model coordinates locomotion and manipulation for unseen object configurations

Deep Dive

Researchers led by Xiaohan Zhang have developed GIRAF, a novel text-conditioned diffusion model that addresses a long-standing challenge in embodied AI: generating coordinated full-body human interactions with articulated objects (e.g., doors, drawers, cabinets). Existing models either focus on static objects and simple activities, or restrict interactions to hand-only manipulation. GIRAF overcomes these limitations by jointly reasoning about locomotion, fine-grained hand-object contact, and object articulation. It introduces three core innovations: an object-centric representation that tightly couples hand-object contact with object surfaces, a mixed-domain training strategy that balances locomotion data with interaction data, and a contact-based augmentation scheme to expand training diversity. The model is designed to generalize across diverse object positions, shapes, and articulation types, producing seamless transitions from approaching an object to manipulating it.

In experiments, GIRAF demonstrated strong generalization to unseen object configurations, surpassing current state-of-the-art methods in both motion realism and task success. The paper has been accepted at the Third Workshop on Human Motion Generation (HuMoGen) at CVPR 2026. This work opens new possibilities for virtual agents that can naturally interact with their environments, robotics training in simulation, and more immersive human-computer interaction. By enabling full-body motion synthesis that adapts to different articulated objects, GIRAF represents a significant step toward generalizable human interaction models for embodied AI systems.

Key Points
  • GIRAF is a text-conditioned diffusion model that generates full-body human interactions with articulated objects like doors and drawers.
  • It uses three innovations: object-centric representation for hand-object contact, mixed-domain training (locomotion + interaction), and contact-based augmentation.
  • The model outperforms prior state-of-the-art methods on unseen object configurations, with applications in robotics training and virtual agents.

Why It Matters

Enables more natural virtual agents and robotics training by generalizing human-object interactions across diverse articulated objects.

📬 Get the top 10 AI stories daily