Gemini 2.5:以先进推理、多模态、长上下文和下一代智能体能力突破前沿

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

诺姆·沙泽尔 Noam Shazeer · · 2025-07-07 · arXiv:2507.06261 ↗ · 被引 3746

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

在本报告中,我们介绍了 Gemini 2.X 模型系列:Gemini 2.5 Pro 和 Gemini 2.5 Flash,以及我们早期的 Gemini 2.0 Flash 和 Flash-Lite 模型。Gemini 2.5 Pro 是我们迄今为止最强大的模型,在前沿编码和推理基准测试中取得了最先进的性能。除了令人难以置信的编码和推理能力外,Gemini 2.5 Pro 还是一款思考型模型,擅长多模态理解,现在能够处理长达 3 小时的视频内容。其长上下文、多模态和推理能力的独特组合可以结合,解锁新的智能体工作流。Gemini 2.5 Flash 以极低的计算和延迟需求提供了出色的推理能力,而 Gemini 2.0 Flash 和 Flash-Lite 则以低延迟和低成本提供高性能。总的来说,Gemini 2.X 模型系列覆盖了模型能力与成本之间的完整帕累托前沿,让用户能够探索复杂智能体问题解决的可能性边界。

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 27)

阅读逐段中英对照全文 →