阐明基于扩散的生成模型的设计空间

Elucidating the Design Space of Diffusion-Based Generative Models

泰罗·卡拉斯 Tero Karras · · 2022-06-01 · arXiv:2206.00364 ↗ · 被引 3702

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们认为,当前基于扩散的生成模型的理论与实践不必要地复杂,并试图通过呈现一个明确区分具体设计选择的设计空间来改善这一状况。这使我们能够识别出采样和训练过程中的几处变更,以及对分数网络的预处理。总之,我们的改进在类条件设置下为 CIFAR-10 取得了新的最先进 FID 1.79,在无条件设置下为 1.97,并且采样速度远快于先前设计(每张图像仅需 35 次网络评估)。为进一步展示其模块化特性,我们表明我们的设计变更显著提高了先前工作中预训练分数网络的效率和质量,包括将先前训练的 ImageNet-64 模型的 FID 从 2.07 提升至接近最先进的 1.55,并在使用我们提出的改进重新训练后达到新的最先进水平 1.36。

We argue that the theory and practice of diffusion-based generative models are currently unnecessarily convoluted and seek to remedy the situation by presenting a design space that clearly separates the concrete design choices. This lets us identify several changes to both the sampling and training processes, as well as preconditioning of the score networks. Together, our improvements yield new state-of-the-art FID of 1.79 for CIFAR-10 in a class-conditional setting and 1.97 in an unconditional setting, with much faster sampling (35 network evaluations per image) than prior designs. To further demonstrate their modular nature, we show that our design changes dramatically improve both the efficiency and quality obtainable with pre-trained score networks from previous work, including improving the FID of a previously trained ImageNet-64 model from 2.07 to near-SOTA 1.55, and after re-training with our proposed improvements to a new SOTA of 1.36.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 10)

阅读逐段中英对照全文 →