Wan:开放且先进的大规模视频生成模型

Wan: Open and Advanced Large-Scale Video Generative Models

通义万相团队 Tongyi Wan Team · · 2025-03-26 · arXiv:2503.20314 ↗ · 被引 2243

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本报告介绍了 Wan,一个全面且开放的视频基础模型套件,旨在推动视频生成的边界。基于主流的扩散 Transformer 范式,Wan 通过一系列创新在生成能力上取得了显著进展,包括我们新颖的 VAE、可扩展的预训练策略、大规模数据整理和自动化评估指标。这些贡献共同提升了模型的性能和通用性。具体而言,Wan 具有四个关键特征:领先性能:Wan 的 14B 模型在海量图像和视频数据集上训练,展示了视频生成在数据和模型规模方面的扩展规律。它在多个内部和外部基准测试中持续超越现有的开源模型以及最先进的商业解决方案,展现出清晰而显著的性能优势。全面性:Wan 提供了两个能力强大的模型,即 1.3B 和 14B 参数,分别用于效率和效果。它还覆盖了多个下游应用,包括图像到视频、指令引导的视频编辑和个性化视频生成,共涉及多达八项任务。消费级效率:1.3B 模型展现出卓越的资源效率,仅需 8.19 GB 显存,使其兼容多种消费级 GPU。开放性:我们开源了 Wan 的整个系列,包括源代码和所有模型,旨在促进视频生成社区的发展。这种开放性力求显著拓展行业中视频制作的创意可能性,并为学术界提供高质量的视频基础模型。所有代码和模型均可访问 https://github.com/Wan-Video/Wan2.1。

This report presents Wan, a comprehensive and open suite of video foundation models designed to push the boundaries of video generation. Built upon the mainstream diffusion transformer paradigm, Wan achieves significant advancements in generative capabilities through a series of innovations, including our novel VAE, scalable pre-training strategies, large-scale data curation, and automated evaluation metrics. These contributions collectively enhance the model's performance and versatility. Specifically, Wan is characterized by four key features: Leading Performance: The 14B model of Wan, trained on a vast dataset comprising billions of images and videos, demonstrates the scaling laws of video generation with respect to both data and model size. It consistently outperforms the existing open-source models as well as state-of-the-art commercial solutions across multiple internal and external benchmarks, demonstrating a clear and significant performance superiority. Comprehensiveness: Wan offers two capable models, i.e., 1.3B and 14B parameters, for efficiency and effectiveness respectively. It also covers multiple downstream applications, including image-to-video, instruction-guided video editing, and personal video generation, encompassing up to eight tasks. Consumer-Grade Efficiency: The 1.3B model demonstrates exceptional resource efficiency, requiring only 8.19 GB VRAM, making it compatible with a wide range of consumer-grade GPUs. Openness: We open-source the entire series of Wan, including source code and all models, with the goal of fostering the growth of the video generation community. This openness seeks to significantly expand the creative possibilities of video production in the industry and provide academia with high-quality video foundation models. All the code and models are available at https://github.com/Wan-Video/Wan2.1.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 23)

阅读逐段中英对照全文 →