Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hunyuan3D-Buffalo 1.0: Unified Multimodal Model for Scalable 3D Generation and Editing

AI By Crimson AI Hugging Face Papers 5 August 2026 · 00:00 11 views
Share: X Telegram

Tencent's Hunyuan3D-Buffalo 1.0 introduces a unified framework for 3D understanding, generation, editing, and part generation, trained on an 87M-scale multimodal corpus.

Hunyuan3D-Buffalo 1.0: Unified Multimodal Model for Scalable 3D Generation and Editing

Key points

Researchers from Tencent have unveiled Hunyuan3D-Buffalo 1.0, a unified multimodal model that tackles multiple 3D tasks within a single architecture. The framework supports 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation, aiming to overcome the scarcity of large-scale 3D multimodal data.

To enable scalable training, the team constructed an 87M-scale 3D multimodal corpus, comprising 25 million understanding samples, 50 million text-to-3D pairs, and 12 million editing pairs generated using Nano3D-v2. This dataset addresses the lack of geometrically consistent editing data, a key bottleneck in unified 3D modeling.

Architecturally, the model combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions.

Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. The analysis further demonstrates that both generation and understanding improve editing, validating the effectiveness of unified 3D multimodal training.

Dataset ComponentNumber of Samples
Understanding samples25M
Text-to-3D pairs50M
Editing pairs12M
Total87M
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1