HunyuanImage 3.0 技术报告

HunyuanImage 3.0 Technical Report

孙兴武 Xingwu Sun · Tencent Hunyuan · 2025-09-28 · arXiv:2509.23951 ↗ · 被引 106

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了 HunyuanImage 3.0,一个原生多模态模型,在自回归框架内统一了多模态理解与生成,其图像生成模块已公开。HunyuanImage 3.0 的成就依赖于几个关键组件,包括精细的数据整理、先进的架构设计、原生思维链方案、渐进式模型预训练、激进式模型后训练,以及支持大规模训练和推理的高效基础设施。凭借这些进步,我们成功训练了一个总参数量超过 800 亿的混合专家(MoE)模型,推理时每个 token 激活 130 亿参数,使其成为迄今为止最大、最强大的开源图像生成模型。我们进行了大量实验,文本-图像对齐和视觉质量的自动与人工评估结果表明,HunyuanImage 3.0 可与之前的最先进模型相媲美。通过发布 HunyuanImage 3.0 的代码和权重,我们希望让社区能够利用最先进的基础模型探索新想法,培育一个充满活力的多模态生态系统。所有开源资源可在 https://github.com/Tencent-Hunyuan/HunyuanImage-3.0 获取。

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module publicly available. The achievement of HunyuanImage 3.0 relies on several key components, including meticulous data curation, advanced architecture design, a native Chain-of-Thoughts schema, progressive model pre-training, aggressive model post-training, and an efficient infrastructure that enables large-scale training and inference. With these advancements, we successfully trained a Mixture-of-Experts (MoE) model comprising over 80 billion parameters in total, with 13 billion parameters activated per token during inference, making it the largest and most powerful open-source image generative model to date. We conducted extensive experiments and the results of automatic and human evaluation of text-image alignment and visual quality demonstrate that HunyuanImage 3.0 rivals previous state-of-the-art models. By releasing the code and weights of HunyuanImage 3.0, we aim to enable the community to explore new ideas with a state-of-the-art foundation model, fostering a dynamic and vibrant multimodal ecosystem. All open source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanImage-3.0

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

阅读逐段中英对照全文 →