In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.
核心贡献 · Key contributions
提出 Gemini 2.X 模型系列,以原生多模态和工具使用能力覆盖模型能力与成本之间的完整帕累托前沿。 Introduces the Gemini 2.X model family, spanning the full Pareto frontier of model capability versus cost with natively multimodal and tool-using models.
Gemini 2.5 Pro 是一款思考模型,在编程、推理、多模态理解以及长达三小时视频的上下文支持方面达到当前最优水平。 Gemini 2.5 Pro is a thinking model with state-of-the-art coding, reasoning, multimodal understanding, and context support for up to three-hour videos.
Gemini 2.5 Flash 提供可控的思考预算,在质量、延迟和成本之间取得平衡,并在多项基准上超越 Gemini 1.5 Pro。 Gemini 2.5 Flash provides a controllable thinking budget, balancing quality, latency, and cost while outperforming Gemini 1.5 Pro on several benchmarks.
训练基础设施的进展包括 TPUv5p 规模扩展、切片级弹性、分阶段静默数据损坏检测,以及基于可验证奖励的强化学习。 Training infrastructure advances include TPUv5p scaling, slice-granularity elasticity, split-phase silent data corruption detection, and reinforcement learning with verifiable rewards.
Gemini Plays Pokémon 展示了智能体式长时程推理能力,利用长上下文工具在 406.5 小时内自主完成游戏。 Agentic long-horizon reasoning is demonstrated by Gemini Plays Pokémon, completing the game autonomously in 406.5 hours using long-context tools.
安全工作通过自动化红队测试和针对间接提示注入攻击的对抗训练,在降低拒答率的同时提升有用性和语气。 Safety work improves helpfulness and tone while reducing refusals, using automated red teaming and adversarial training against indirect prompt injection attacks.
局限 · Limitations
Gemini 2.5 Pro 难以直接处理原始游戏屏幕像素,需要借助 RAM 到文本的转换;在 Pokémon 智能体中,移除视觉输入几乎不影响性能。 Gemini 2.5 Pro struggled with raw game-screen pixels, requiring RAM-to-text translation; ablating visual input barely reduced performance in the Pokémon agent.
在上下文超过约 10 万 token 后,智能体倾向于重复历史动作而非综合新计划,暴露出检索与生成式推理之间的差距。 Beyond roughly 100k-token contexts, the agent favored repeating historical actions over synthesizing new plans, exposing a gap between retrieval and generative reasoning.
尽管对提示注入攻击的抵抗力有所提升,Gemini 2.5 Pro 仍不如 Flash 稳健,表明更强的能力会限制安全缓解措施的效果。 Despite improved resilience against prompt-injection attacks, Gemini 2.5 Pro remains less robust than Flash, suggesting higher capability constrains security mitigations.
图像到文本的政策违规率略有上升,人工审查显示损失主要集中在露骨色情或仇恨类创意内容上。 Image-to-text policy violations slightly increased, with manual review showing losses concentrated around explicit sexual or hateful creative content.
Gemini 2.5 Pro 尚未达到关键能力等级,但部分能力的显著提升仍需持续监测,以防范未来的危险能力风险。 Gemini 2.5 Pro did not reach critical capability levels, but notable capability increases warrant continued monitoring for future dangerous-capability risks.