视频扩散模型

Video Diffusion Models

乔纳森·何 Jonathan Ho · Google · 2022-04-07 · arXiv:2204.03458 ↗ · 被引 2666

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

生成时间上连贯的高保真视频是生成建模研究中的一个重要里程碑。我们通过提出一种用于视频生成的扩散模型,朝着这一里程碑取得了进展,该模型展示了非常有前景的初步结果。我们的模型是标准图像扩散架构的自然扩展,它支持从图像和视频数据中联合训练,我们发现这可以减少小批量梯度的方差并加速优化。为了生成长时间且分辨率更高的视频,我们引入了一种新的条件采样技术,用于空间和时间视频扩展,其性能优于先前提出的方法。我们首次展示了在大规模文本条件视频生成任务上的结果,以及在视频预测和无条件视频生成的既定基准上达到的最先进结果。补充材料可在 https://video-diffusion.github.io/ 获取。

Generating temporally coherent high fidelity video is an important milestone in generative modeling research. We make progress towards this milestone by proposing a diffusion model for video generation that shows very promising initial results. Our model is a natural extension of the standard image diffusion architecture, and it enables jointly training from image and video data, which we find to reduce the variance of minibatch gradients and speed up optimization. To generate long and higher resolution videos we introduce a new conditional sampling technique for spatial and temporal video extension that performs better than previously proposed methods. We present the first results on a large text-conditioned video generation task, as well as state-of-the-art results on established benchmarks for video prediction and unconditional video generation. Supplementary material is available at https://video-diffusion.github.io/

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 11)

阅读逐段中英对照全文 →