具有深度语言理解能力的逼真文本到图像扩散模型

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

奇特万·萨哈里亚 Chitwan Saharia · Google · 2022-05-23 · arXiv:2205.11487 ↗ · 被引 8550

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了 Imagen,一种具有前所未有的逼真度和深度语言理解能力的文本到图像扩散模型。Imagen 建立在大型 Transformer 语言模型理解文本的能力以及扩散模型高保真图像生成的优势之上。我们的关键发现是,仅在文本语料上预训练的通用大型语言模型(如 T5)在编码文本用于图像合成方面出奇地有效:增加 Imagen 中语言模型的规模比增加图像扩散模型的规模更能提升样本保真度和图像-文本对齐。Imagen 在 COCO 数据集上取得了 7.27 的 FID 分数,达到了新的最先进水平,且从未在 COCO 上训练过;人类评估者认为 Imagen 样本在图像-文本对齐方面与 COCO 数据本身相当。为了更深入地评估文本到图像模型,我们引入了 DrawBench,一个全面且具有挑战性的基准。通过 DrawBench,我们将 Imagen 与近期方法(包括 VQ-GAN+CLIP、潜在扩散模型和 DALL-E 2)进行比较,发现人类评估者在样本质量和图像-文本对齐方面都更偏好 Imagen。

We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image generation. Our key discovery is that generic large language models (e.g. T5), pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: increasing the size of the language model in Imagen boosts both sample fidelity and image-text alignment much more than increasing the size of the image diffusion model. Imagen achieves a new state-of-the-art FID score of 7.27 on the COCO dataset, without ever training on COCO, and human raters find Imagen samples to be on par with the COCO data itself in image-text alignment. To assess text-to-image models in greater depth, we introduce DrawBench, a comprehensive and challenging benchmark for text-to-image models. With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment. See https://imagen.research.google/ for an overview of the results.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 15)

阅读逐段中英对照全文 →