Gemini 1.5:解锁跨数百万 token 上下文的 multimodal 理解

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

杰米斯·哈萨比斯 Demis Hassabis · · 2024-03-08 · arXiv:2403.05530 ↗ · 被引 3804

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本报告介绍了 Gemini 1.5 系列模型,它们是下一代计算高效的多模态模型,能够从数百万 token 的上下文中回忆和推理细粒度信息,包括多个长文档以及数小时的视频和音频。该系列包含两个新模型:(1)更新的 Gemini 1.5 Pro,在绝大多数能力和基准测试上超越了 2 月版本;(2)Gemini 1.5 Flash,一种更轻量化的变体,专为效率设计而质量下降极小。Gemini 1.5 模型在跨模态的长上下文检索任务上实现了近乎完美的召回,在长文档问答、长视频问答和长上下文自动语音识别方面提升了最先进水平,并在广泛的基准测试中匹配或超越了 Gemini 1.0 Ultra 的最先进性能。研究 Gemini 1.5 的长上下文能力极限时,我们发现其在至少 1000 万 token 上持续改进下一个 token 预测并实现近乎完美的检索(>99%),这是对现有模型(如 Claude 3.0 的 200k 和 GPT-4 Turbo 的 128k)的代际飞跃。最后,我们强调实际用例,例如 Gemini 1.5 与专业人员合作完成任务,在 10 个不同工作类别中实现 26%至 75%的时间节省,以及大语言模型在前沿领域令人惊喜的新能力:给定一本 Kalamang 语语法手册(全球使用该语言者不足 200 人),模型学习将英语翻译成 Kalamang,其水平与从相同内容学习的人相当。

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February version on the great majority of capabilities and benchmarks; (2) Gemini 1.5 Flash, a more lightweight variant designed for efficiency with minimal regression in quality. Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks. Studying the limits of Gemini 1.5's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens, a generational leap over existing models such as Claude 3.0 (200k) and GPT-4 Turbo (128k). Finally, we highlight real-world use cases, such as Gemini 1.5 collaborating with professionals on completing their tasks achieving 26 to 75% time savings across 10 different job categories, as well as surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person who learned from the same content.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 64)

阅读逐段中英对照全文 →