Research & Papers

Hunyuan3D-Buffalo 1.0 unifies 3D generation and editing with 87M-scale dataset

One model handles text-to-3D, editing, and part generation — trained on 87M multimodal samples

Deep Dive

Researchers from Tencent have released Hunyuan3D-Buffalo 1.0, a unified multimodal model that combines 3D understanding, generation, and editing into a single architecture. Unlike previous approaches that treat these tasks separately, the model integrates a Vision-Language Model (Hunyuan3D-VLM) for semantic, structural, and spatial reasoning with a diffusion transformer (Hunyuan3D DiT) for high-fidelity 3D synthesis. This design lets the VLM provide multimodal semantic conditions for generation, while editing and part-generation tasks additionally condition the diffusion process on the source object representation to preserve unedited regions.

To train the model at scale, the team built an 87M-sample 3D multimodal corpus — the largest of its kind — comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs synthesized using their Nano3D-v2 pipeline. This data diversity addresses the historical scarcity of geometrically consistent 3D editing data. On benchmarks, Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance in text-to-3D generation and 3D editing, while also showing strong understanding and part-generation capabilities. Notably, the researchers found that joint training on generation and understanding improves editing quality, validating the unified approach for scalable 3D multimodal learning.

Key Points
  • Hunyuan3D-Buffalo 1.0 combines Hunyuan3D-VLM with a 3D DiT for unified understanding, generation, editing, and part generation.
  • Trained on an 87M-scale 3D corpus: 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated via Nano3D-v2.
  • Achieves state-of-the-art performance on text-to-3D and 3D editing benchmarks, with joint training improving editing quality.

Why It Matters

Unified 3D models cut pipeline complexity, enabling consistent text-to-3D, editing, and part generation for games, VFX, and design workflows.

📬 Get the top 10 AI stories daily