Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.
核心贡献 · Key contributions
提出 Seedance 2.0,一个原生多模态音视频联合生成模型,采用统一大规模架构,支持文本、图像、音频和视频输入。 Proposes Seedance 2.0, a native multimodal audio-video joint generation model with unified large-scale architecture supporting text, image, audio, and video inputs.
在文生视频、图生视频和参考生视频任务上取得最优结果,在 Arena 排行榜和 SeedVideoBench 2.0 的所有评估维度上均排名第一。 Achieves state-of-the-art results on T2V, I2V, and R2V tasks, ranking first on Arena leaderboards and SeedVideoBench 2.0 across all evaluated dimensions.
引入全面的多模态参考与编辑能力,涵盖主体、运动、风格、特效控制以及视频续写/扩展,适用于多种创作流程。 Introduces comprehensive multimodal reference and editing capabilities—subject, motion, style, visual-effects control, and video continuation/extension—for diverse creative workflows.
生成高保真同步双耳音频,支持双声道输出、精确的视听对齐和强大的音频指令遵循。 Generates high-fidelity synchronized binaural audio with dual-channel output, precise audio-visual alignment, and strong audio instruction following.
改进真实世界复杂性建模,尤其是人体运动自然度、物理合理性和多主体交互稳定性,缓解常见生成伪影。 Improves real-world complexity modeling—human motion naturalness, physical plausibility, and multi-subject interaction stability—mitigating common artifacts.
局限 · Limitations
作者承认 Seedance 2.0 仍不完美,生成输出仍有改进空间。 The authors acknowledge Seedance 2.0 remains imperfect, with generation outputs still having room for improvement.
残余问题包括轻微形变伪影、边缘场景运动合理性不足、高频视觉噪声、音频失真以及多说话人场景中的口型同步误差。 Residual issues include minor deformation artifacts, motion plausibility in edge cases, high-frequency visual noise, audio distortion, and lip-sync errors in multi-speaker scenes.
多主体一致性、文字还原准确率以及复杂编辑任务的表现仍需优化。 Multi-subject consistency, text restoration accuracy, and complex editing task performance still require optimization.
视频扩展质量落后于 Veo 3.1;由于生成更动态,首帧保真度低于 Kling 3 Omni。 Video extension quality trails Veo 3.1, and first-frame preservation is lower than Kling 3 Omni due to more dynamic generation.
图像-音频联合参考条件化仍然困难,原生输出仅支持 4-15 秒、480p/720p 分辨率。 Joint image-audio reference conditioning remains difficult, with native output limited to 4–15 seconds at 480p/720p resolution.