本报告介绍了 Hunyuan3D 2.5,这是一个强大的 3D 扩散模型套件,旨在生成高保真且细节丰富的纹理 3D 资产。Hunyuan3D 2.5 沿用了其前身 Hunyuan3D 2.0 的两阶段流水线,同时在形状和纹理生成方面展示了显著进步。在形状生成方面,我们引入了一个新的形状基础模型——LATTICE,该模型通过大规模高质量数据集、模型规模和计算量进行训练。我们最大的模型达到了 10B 参数,能够生成清晰且细节丰富的 3D 形状,具有精确的图像-3D 跟随能力,同时保持网格表面干净光滑,显著缩小了生成形状与手工制作形状之间的差距。在纹理生成方面,通过从 Hunyuan3D 2.0 Paint 模型扩展而来的新颖多视图架构,升级了基于物理的渲染(PBR)。我们的广泛评估表明,Hunyuan3D 2.5 在形状和端到端纹理生成方面均显著优于以往方法。
In this report, we present Hunyuan3D 2.5, a robust suite of 3D diffusion models aimed at generating high-fidelity and detailed textured 3D assets. Hunyuan3D 2.5 follows two-stages pipeline of its previous version Hunyuan3D 2.0, while demonstrating substantial advancements in both shape and texture generation. In terms of shape generation, we introduce a new shape foundation model -- LATTICE, which is trained with scaled high-quality datasets, model-size, and compute. Our largest model reaches 10B parameters and generates sharp and detailed 3D shape with precise image-3D following while keeping mesh surface clean and smooth, significantly closing the gap between generated and handcrafted 3D shapes. In terms of texture generation, it is upgraded with phyiscal-based rendering (PBR) via a novel multi-view architecture extended from Hunyuan3D 2.0 Paint model. Our extensive evaluation shows that Hunyuan3D 2.5 significantly outperforms previous methods in both shape and end-to-end texture generation.
核心贡献 · Key contributions
提出 LATTICE,一个基于高质量数据规模扩张训练的 100 亿参数形状基础模型。 Introduces LATTICE, a 10B-parameter shape foundation model trained on scaled high-quality data.
生成锐利、细节丰富的 3D 形状,实现精确的图像-3D 对齐和光滑表面。 Achieves sharp, detailed 3D shapes with precise image-3D alignment and smooth surfaces.
通过新颖的多视图架构和双通道注意力机制,将纹理生成扩展到 PBR 材质。 Extends texture generation to PBR materials via a novel multi-view architecture with dual-channel attention.
提出双阶段分辨率增强策略,改善纹理与几何的对齐质量。 Proposes a dual-phase resolution enhancement strategy for better texture-geometry alignment.
在形状和纹理生成上超越最先进的开源和商业模型。 Outperforms state-of-the-art open-source and commercial models in shape and texture generation.
展示了模型规模扩张带来的稳定提升,缩小了与手工 3D 资产的差距。 Demonstrates stable improvement with model scaling, closing the gap to handcrafted 3D assets.
局限 · Limitations
ULIP 和 Uni3D 等评估指标可能无法完全反映视觉质量的提升。 Evaluation metrics like ULIP and Uni3D may not fully capture visual quality improvements.
高分辨率多视图训练需要大量内存,限制了视图数量。 High-resolution multi-view training requires substantial memory, limiting view count.
PBR 材质生成在从反照率中解耦光照方面仍面临挑战。 PBR material generation still faces challenges in decoupling illumination from albedo.
模型在极其复杂或罕见物体类别上的性能未经过广泛测试。 The model's performance on extremely complex or rare object categories is not extensively tested.
由于多步采样和高分辨率处理,推理速度可能受限。 Inference speed may be limited due to multi-step sampling and high-resolution processing.