通过估计数据分布的梯度进行生成建模

Generative Modeling by Estimating Gradients of the Data Distribution

宋飏 Yang Song · Stanford · 2019-07-12 · arXiv:1907.05600 ↗ · 被引 5692

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们引入了一种新的生成模型,其中样本通过朗之万动力学生成,利用分数匹配估计的数据分布梯度。由于当数据位于低维流形上时,梯度可能定义不良且难以估计,我们使用不同级别的高斯噪声扰动数据,并联合估计相应的分数,即所有噪声级别下扰动数据分布的梯度向量场。对于采样,我们提出了一种退火朗之万动力学,其中随着采样过程接近数据流形,我们使用对应逐渐减小噪声级别的梯度。我们的框架允许灵活的模型架构,训练期间无需采样或使用对抗方法,并提供了一个可用于原则性模型比较的学习目标。我们的模型在 MNIST、CelebA 和 CIFAR-10 数据集上生成的样本可与 GAN 相媲美,在 CIFAR-10 上实现了 8.87 的新最优初始分数。此外,我们通过图像修复实验证明了我们的模型学习了有效的表示。

We introduce a new generative model where samples are produced via Langevin dynamics using gradients of the data distribution estimated with score matching. Because gradients can be ill-defined and hard to estimate when the data resides on low-dimensional manifolds, we perturb the data with different levels of Gaussian noise, and jointly estimate the corresponding scores, i.e., the vector fields of gradients of the perturbed data distribution for all noise levels. For sampling, we propose an annealed Langevin dynamics where we use gradients corresponding to gradually decreasing noise levels as the sampling process gets closer to the data manifold. Our framework allows flexible model architectures, requires no sampling during training or the use of adversarial methods, and provides a learning objective that can be used for principled model comparisons. Our models produce samples comparable to GANs on MNIST, CelebA and CIFAR-10 datasets, achieving a new state-of-the-art inception score of 8.87 on CIFAR-10. Additionally, we demonstrate that our models learn effective representations via image inpainting experiments.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →