SIMA 2:面向虚拟世界的通用具身智能体

SIMA 2: A Generalist Embodied Agent for Virtual Worlds

杰米斯·哈萨比斯 Demis Hassabis · Google DeepMind · 2025-12-04 · arXiv:2512.04797 ↗ · 被引 14

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 SIMA 2,这是一个通用具身智能体,能够理解并在多种 3D 虚拟世界中行动。它基于 Gemini 基础模型构建,代表了在具身环境中进行主动、目标导向交互的重要一步。与先前局限于简单语言指令的工作(如 SIMA 1)不同,SIMA 2 作为一个交互式伙伴,能够推理高层目标、与用户对话,并处理通过语言和图像给出的复杂指令。在多样化的游戏组合中,SIMA 2 大幅缩小了与人类表现的差距,并展现出对未见环境的强大泛化能力,同时保留了基础模型的核心推理能力。此外,我们展示了开放式自我改进的能力:通过利用 Gemini 生成任务并提供奖励,SIMA 2 可以在新环境中从零开始自主学习新技能。这项工作验证了一条通往创建多功能且持续学习智能体的路径,这些智能体最终将适用于虚拟世界乃至物理世界。

We introduce SIMA 2, a generalist embodied agent that understands and acts in a wide variety of 3D virtual worlds. Built upon a Gemini foundation model, SIMA 2 represents a significant step toward active, goal-directed interaction within an embodied environment. Unlike prior work (e.g., SIMA 1) limited to simple language commands, SIMA 2 acts as an interactive partner, capable of reasoning about high-level goals, conversing with the user, and handling complex instructions given through language and images. Across a diverse portfolio of games, SIMA 2 substantially closes the gap with human performance and demonstrates robust generalization to previously unseen environments, all while retaining the base model's core reasoning capabilities. Furthermore, we demonstrate a capacity for open-ended self-improvement: by leveraging Gemini to generate tasks and provide rewards, SIMA 2 can autonomously learn new skills from scratch in a new environment. This work validates a path toward creating versatile and continuously learning agents for both virtual and, eventually, physical worlds.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

阅读逐段中英对照全文 →