Gemini:一个高度能干的多模态模型家族

Gemini: A Family of Highly Capable Multimodal Models

诺姆·沙泽尔 Noam Shazeer · · 2023-12-19 · arXiv:2312.11805 ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本报告介绍了一个新的多模态模型家族 Gemini,它在图像、音频、视频和文本理解方面表现出非凡的能力。Gemini 家族包括 Ultra、Pro 和 Nano 三种尺寸,适用于从复杂推理任务到内存受限的端侧用例等多种应用。在广泛的基准测试评估中,我们最强大的 Gemini Ultra 模型在 32 项基准测试中的 30 项上取得了最优性能——值得注意的是,它是首个在广受研究的考试基准 MMLU 上达到人类专家水平的模型,并且在我们考察的 20 个多模态基准测试中每一项都提升了最优水平。我们相信,Gemini 家族在跨模态推理和语言理解方面的新能力将能够支持多种多样的用例。我们讨论了如何通过包括 Gemini、Gemini Advanced、Google AI Studio 和 Cloud Vertex AI 在内的服务,负责任地对 Gemini 模型进行后训练和部署,并交付给用户。

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultra model advances the state of the art in 30 of 32 of these benchmarks - notably being the first model to achieve human-expert performance on the well-studied exam benchmark MMLU, and improving the state of the art in every one of the 20 multimodal benchmarks we examined. We believe that the new capabilities of the Gemini family in cross-modal reasoning and language understanding will enable a wide variety of use cases. We discuss our approach toward post-training and deploying Gemini models responsibly to users through services including Gemini, Gemini Advanced, Google AI Studio, and Cloud Vertex AI.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 31)

阅读逐段中英对照全文 →