HunyuanVideo 1.5 技术报告

HunyuanVideo 1.5 Technical Report

孙兴武 Xingwu Sun · Tencent Hunyuan · 2025-11-24 · arXiv:2511.18870 ↗ · 被引 78

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了 HunyuanVideo 1.5,一个轻量级但功能强大的开源视频生成模型,仅用 83 亿参数就实现了最先进的视觉质量和运动连贯性,能够在消费级 GPU 上进行高效推理。这一成就建立在几个关键组件之上,包括精心设计的数据整理、采用选择性滑动窗口注意力(SSTA)的高级 DiT 架构、通过字形感知文本编码增强的双语理解、渐进式预训练和后训练,以及高效的视频超分辨率网络。利用这些设计,我们开发了一个统一的框架,能够跨多种时长和分辨率进行高质量的文本到视频和图像到视频生成。大量实验表明,这个紧凑而高效的模型在开源视频生成模型中树立了新的标杆。通过发布代码和模型权重,我们为社区提供了一个高性能基础,降低了视频创作和研究的门槛,使先进的视频生成技术惠及更广泛的受众。所有开源资源均可公开获取:https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5。

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architecture featuring selective and sliding tile attention (SSTA), enhanced bilingual understanding through glyph-aware text encoding, progressive pre-training and post-training, and an efficient video super-resolution network. Leveraging these designs, we developed a unified framework capable of high-quality text-to-video and image-to-video generation across multiple durations and resolutions. Extensive experiments demonstrate that this compact and proficient model establishes a new state-of-the-art among open-source video generation models. By releasing the code and model weights, we provide the community with a high-performance foundation that lowers the barrier to video creation and research, making advanced video generation accessible to a broader audience. All open-source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →