ImageBind:一个嵌入空间绑定所有

ImageBind: One Embedding Space To Bind Them All

罗斯·吉尔希克 Ross Girshick · · 2023-05-09 · arXiv:2305.05665 ↗ · 被引 1660

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出 ImageBind,一种学习跨六种不同模态(图像、文本、音频、深度、热成像和 IMU 数据)的联合嵌入的方法。我们表明,训练这样的联合嵌入并不需要所有配对数据的组合,仅图像配对数据就足以将模态绑定在一起。ImageBind 可以利用最近的大规模视觉-语言模型,并通过与图像的自然配对将其零样本能力扩展到新模态。它实现了开箱即用的新涌现应用,包括跨模态检索、用算术组合模态、跨模态检测和生成。涌现能力随着图像编码器的强度而提高,我们在跨模态的涌现零样本识别任务上达到了新的最先进水平,优于专业监督模型。最后,我们展示了强大的少样本识别结果,优于先前的工作,并且 ImageBind 作为评估视觉模型在视觉和非视觉任务上的新方法。

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint embedding, and only image-paired data is sufficient to bind the modalities together. ImageBind can leverage recent large scale vision-language models, and extends their zero-shot capabilities to new modalities just by using their natural pairing with images. It enables novel emergent applications 'out-of-the-box' including cross-modal retrieval, composing modalities with arithmetic, cross-modal detection and generation. The emergent capabilities improve with the strength of the image encoder and we set a new state-of-the-art on emergent zero-shot recognition tasks across modalities, outperforming specialist supervised models. Finally, we show strong few-shot recognition results outperforming prior work, and that ImageBind serves as a new way to evaluate vision models for visual and non-visual tasks.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →