Seedance 1.0:探索视频生成模型的边界

Seedance 1.0: Exploring the Boundaries of Video Generation Models

高宇 Yu Gao · ByteDance Seed · 2025-06-10 · arXiv:2506.09113 ↗ · 被引 223

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

扩散建模的显著突破推动了视频生成的快速进步,但当前的基础模型在同时平衡提示遵循、运动合理性和视觉质量方面仍面临关键挑战。在本报告中,我们介绍了 Seedance 1.0,这是一个高性能且推理高效的视频基础生成模型,它整合了多项核心技术改进:(i)多源数据整理,辅以精确且有意义的视频字幕,实现了跨不同场景的全面学习;(ii)高效的架构设计与提出的训练范式,原生支持多镜头生成,并联合学习文本到视频和图像到视频任务;(iii)精心优化的训练后方法,利用细粒度监督微调和具有多维奖励机制的视频特定 RLHF,实现全面的性能提升;(iv)通过多阶段蒸馏策略和系统级优化实现约 10 倍的推理加速。Seedance 1.0 仅需 41.4 秒(NVIDIA-L20)即可生成 5 秒 1080p 分辨率的视频。与最先进的视频生成模型相比,Seedance 1.0 以其高质量和快速的视频生成脱颖而出,具有卓越的时空流畅性和结构稳定性,在复杂的多主体上下文中精确遵循指令,以及具有一致主体表示的原生多镜头叙事连贯性。

Notable breakthroughs in diffusion modeling have propelled rapid improvements in video generation, yet current foundational model still face critical challenges in simultaneously balancing prompt following, motion plausibility, and visual quality. In this report, we introduce Seedance 1.0, a high-performance and inference-efficient video foundation generation model that integrates several core technical improvements: (i) multi-source data curation augmented with precision and meaningful video captioning, enabling comprehensive learning across diverse scenarios; (ii) an efficient architecture design with proposed training paradigm, which allows for natively supporting multi-shot generation and jointly learning of both text-to-video and image-to-video tasks. (iii) carefully-optimized post-training approaches leveraging fine-grained supervised fine-tuning, and video-specific RLHF with multi-dimensional reward mechanisms for comprehensive performance improvements; (iv) excellent model acceleration achieving ~10x inference speedup through multi-stage distillation strategies and system-level optimizations. Seedance 1.0 can generate a 5-second video at 1080p resolution only with 41.4 seconds (NVIDIA-L20). Compared to state-of-the-art video generation models, Seedance 1.0 stands out with high-quality and fast video generation having superior spatiotemporal fluidity with structural stability, precise instruction adherence in complex multi-subject contexts, native multi-shot narrative coherence with consistent subject representation.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 26)

阅读逐段中英对照全文 →