生成式多模态模型是上下文学习者

Generative Multimodal Models are In-Context Learners

王鑫龙 Xinlong Wang · BAAI · 2023-12-20 · arXiv:2312.13286 ↗ · 被引 512

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

人类能够轻松地在上下文中(即仅凭少量示例或简单指令)解决多模态任务,而当前的多模态系统在很大程度上难以模仿这一点。在这项工作中,我们证明了通过有效扩展规模,可以显著增强大型多模态模型的与任务无关的上下文学习能力。我们引入了 Emu2,一个拥有 370 亿参数的生成式多模态模型,在大型多模态序列上使用统一的自回归目标进行训练。Emu2 展现出强大的多模态上下文学习能力,甚至能够解决需要即时推理的任务,如视觉提示和基于对象的生成。该模型在少样本设置下的多个多模态理解任务上创下了新纪录。当通过指令微调以遵循特定指令时,Emu2 在具有挑战性的任务上进一步达到了新的最先进水平,例如大型多模态模型的问答基准和开放式的主题驱动生成。这些成就表明,Emu2 可以作为广泛多模态任务的基础模型和通用接口。代码和模型已公开,以促进未来的研究。

The human ability to easily solve multimodal tasks in context (i.e., with only a few demonstrations or simple instructions), is what current multimodal systems have largely struggled to imitate. In this work, we demonstrate that the task-agnostic in-context learning capabilities of large multimodal models can be significantly enhanced by effective scaling-up. We introduce Emu2, a generative multimodal model with 37 billion parameters, trained on large-scale multimodal sequences with a unified autoregressive objective. Emu2 exhibits strong multimodal in-context learning abilities, even emerging to solve tasks that require on-the-fly reasoning, such as visual prompting and object-grounded generation. The model sets a new record on multiple multimodal understanding tasks in few-shot settings. When instruction-tuned to follow specific instructions, Emu2 further achieves new state-of-the-art on challenging tasks such as question answering benchmarks for large multimodal models and open-ended subject-driven generation. These achievements demonstrate that Emu2 can serve as a base model and general-purpose interface for a wide range of multimodal tasks. Code and models are publicly available to facilitate future research.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 11)

阅读逐段中英对照全文 →