Gemini Robotics: Bringing AI into the Physical World
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→大型多模态模型的最新进展催生了数字领域显著的通用能力,但将其转化为机器人等物理智能体仍是一项重大挑战。本报告介绍了一个专为机器人设计、基于 Gemini 2.0 构建的新 AI 模型系列。我们推出了 Gemini Robotics,一种先进的视觉-语言-动作(VLA)通用模型,能够直接控制机器人。Gemini Robotics 执行流畅且反应灵敏的动作,以处理各种复杂的操作任务,同时对物体类型和位置的变化具有鲁棒性,能应对未见过的环境,并遵循多样化的开放词汇指令。我们表明,通过额外的微调,Gemini Robotics 可以专门化以获得新能力,包括解决长时域、高灵巧性任务,从仅 100 次演示中学习新的短时域任务,以及适应全新的机器人形态。这得益于 Gemini Robotics 建立在 Gemini Robotics-ER 模型之上,这是我们在本工作中引入的第二个模型。Gemini Robotics-ER(具身推理)将 Gemini 的多模态推理能力扩展到物理世界,增强了空间和时间理解。这实现了与机器人相关的能力,包括物体检测、指向、轨迹和抓取预测,以及多视图对应和 3D 边界框预测。我们展示了这种新颖组合如何支持各种机器人应用。我们还讨论并解决了与这类新型机器人基础模型相关的重要安全问题。Gemini Robotics 系列标志着朝着开发通用机器人迈出的重要一步,实现了 AI 在物理世界中的潜力。
Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report introduces a new family of AI models purposefully designed for robotics and built upon the foundation of Gemini 2.0. We present Gemini Robotics, an advanced Vision-Language-Action (VLA) generalist model capable of directly controlling robots. Gemini Robotics executes smooth and reactive movements to tackle a wide range of complex manipulation tasks while also being robust to variations in object types and positions, handling unseen environments as well as following diverse, open vocabulary instructions. We show that with additional fine-tuning, Gemini Robotics can be specialized to new capabilities including solving long-horizon, highly dexterous tasks, learning new short-horizon tasks from as few as 100 demonstrations and adapting to completely novel robot embodiments. This is made possible because Gemini Robotics builds on top of the Gemini Robotics-ER model, the second model we introduce in this work. Gemini Robotics-ER (Embodied Reasoning) extends Gemini's multimodal reasoning capabilities into the physical world, with enhanced spatial and temporal understanding. This enables capabilities relevant to robotics including object detection, pointing, trajectory and grasp prediction, as well as multi-view correspondence and 3D bounding box predictions. We show how this novel combination can support a variety of robotics applications. We also discuss and address important safety considerations related to this new class of robotics foundation models. The Gemini Robotics family marks a substantial step towards developing general-purpose robots that realizes AI's potential in the physical world.
近期大型多模态模型的进展催生了数字领域非凡的通用能力,但这些能力向机器人等物理智能体的转化仍是一个重大挑战。通常有用的机器人需要能够理解其周围的物理世界,并与之进行胜任且安全的交互。本报告介绍了一个专为机器人设计、基于 Gemini 2.0 构建的新 AI 模型系列。我们提出了 Gemini Robotics,一个先进的视觉-语言-动作(VLA)通用模型,能够直接控制机器人。Gemini Robotics 执行平滑且反应灵敏的动作,以处理各种复杂的操作任务,同时对物体类型和位置的变化具有鲁棒性,能处理未见过的环境,并遵循多样化的开放词汇指令。我们表明,通过额外的微调,Gemini Robotics 可以专门化以获得新能力,包括解决长时域、高灵巧度的任务(如折叠折纸狐狸或玩纸牌游戏),从少至 100 次演示中学习新的短时域任务,以及适应全新的机器人形态,包括双臂平台和高自由度人形机器人。这之所以成为可能,是因为 Gemini Robotics 构建于 Gemini Robotics-ER 模型之上,这是我们在本工作中引入的第二个模型。Gemini Robotics-ER(具身推理)将 Gemini 的多模态推理能力扩展到物理世界,增强了空间和时间理解。这实现了与机器人相关的能力,包括物体检测、指向、轨迹和抓取预测,以及多视角对应和 3D 边界框预测形式的 3D 理解。我们展示了这种新颖组合如何支持各种机器人应用,例如零样本(通过机器人代码生成)或少样本(通过上下文学习)。我们还讨论并解决了与这类新型机器人基础模型相关的重要安全问题。Gemini Robotics 系列标志着朝着开发通用机器人、实现 AI 在物理世界中的潜力迈出了重要一步。
Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. Generally useful robots need to be able to make sense of the physical world around them, and interact with it competently and safely. This report introduces a new family of AI models purposefully designed for robotics and built upon the foundation of Gemini 2.0. We present Gemini Robotics, an advanced Vision-Language-Action (VLA) generalist model capable of directly controlling robots. Gemini Robotics executes smooth and reactive movements to tackle a wide range of complex manipulation tasks while also being robust to variations in object types and positions, handling unseen environments as well as following diverse, open vocabulary instructions. We show that with additional fine-tuning, Gemini Robotics can be specialized to new capabilities including solving long-horizon, highly dexterous tasks like folding an origami fox or playing a game of cards, learning new short-horizon tasks from as few as 100 demonstrations, adapting to completely novel robot embodiments including a bi-arm platform and a high degrees-of-freedom humanoid. This is made possible because Gemini Robotics builds on top of the Gemini Robotics-ER model, the second model we introduce in this work. Gemini Robotics-ER (Embodied Reasoning) extends Gemini’s multimodal reasoning capabilities into the physical world, with enhanced spatial and temporal understanding. This enables capabilities relevant to robotics including object detection, pointing, trajectory and grasp prediction, as well as 3D understanding in the form of multi-view correspondence and 3D bounding box predictions. We show how this novel combination can support a variety of robotics applications, e.g., zero-shot (via robot code generation), or few-shot (via in-context learning). We also discuss and address important safety considerations related to this new class of robotics foundation models. The Gemini Robotics family marks a substantial step towards developing general-purpose robots that realize AI’s potential in the physical world.
现代人工智能(AI)模型在大规模数据集上进行预训练,取得了显著进展,重新定义了信息处理方式,在文本、图像、音频和视频等多种模态上展示了熟练性和泛化能力。这为数字领域内的交互式和辅助系统开辟了广阔的前景,从多模态聊天机器人到虚拟助手不一而足。然而,要在物理世界中实现通用自主 AI 的潜力,需要从数字世界进行实质性转变,物理世界中的 AI 智能体必须展现出强大的人类级具身推理能力:即包含对在固有物理具身世界中操作和行动至关重要的基本概念的世界知识。虽然作为人类,我们理所当然地认为具身推理能力(如感知环境的 3D 结构、解释复杂的物体间关系或理解直观物理)是理所当然的,但这些能力构成了任何具身 AI 智能体的重要基础。此外,具身 AI 智能体还必须超越被动理解真实世界的空间和物理概念;它还必须学会采取对外部环境产生直接影响的行为,弥合被动感知与主动物理交互之间的鸿沟。
The remarkable progress of modern artificial intelligence (AI) models – with pre-training on large-scale datasets – has redefined information processing, demonstrating proficiency and generalization across diverse modalities such as text, images, audio, and video. This has opened a vast landscape of opportunities for interactive and assistive systems within the digital realm, ranging from multimodal chatbots to virtual assistants. However, realizing the potential of general-purpose autonomous AI in the physical world requires a substantial shift from the digital world, where physically grounded AI agents must demonstrate robust human-level embodied reasoning: The set of world knowledge that encompasses the fundamental concepts which are critical for operating and acting in an inherently physically embodied world. While, as humans, we take for granted our embodied reasoning abilities – such as perceiving the 3D structure of environments, interpreting complex inter-object relationships, or understanding intuitive physics – these capabilities form an important basis for any embodied AI agent. Furthermore, an embodied AI agent must also go beyond passively understanding the spatial and physical concepts of the real world; it must also learn to take actions that have direct effects on their external environment, bridging the gap between passive perception and active physical interaction.
随着机器人硬件的最新进展,创建能够执行高度灵巧任务的具身 AI 智能体具有令人兴奋的潜力。考虑到这一点,我们不禁要问:如何赋予最先进的数字 AI 模型以具身推理能力,使其能够以通用且灵巧的方式与我们的世界互动?
With the recent advancements in robotics hardware, there is exciting potential for creating embodied AI agents that can perform highly dexterous tasks. With this in mind, we ask: What would it take to endow a state-of-the-art digital AI model with the embodied reasoning capabilities needed to interact with our world in a general and dexterous manner?
我们的论点基于利用前沿视觉-语言模型(VLM)中固有的高级多模态理解和推理能力,例如 Gemini 2.0。这些基础模型提供的通用理解能力,以及它们解释视觉输入和复杂文本指令的能力,为构建具身智能体奠定了坚实的基础。这一努力依赖于两个基本组成部分。首先,Gemini 需要获得强大的具身推理能力,能够理解物理世界丰富的几何和时空细节。其次,我们必须通过使 Gemini 能够说物理行动的语言,理解接触物理、动力学以及现实世界交互的复杂性,将这种具身推理扎根于物理世界。最终,这些要素必须结合起来,以实现对真实世界机器人的快速、安全和灵巧控制。
Our thesis is predicated on harnessing the advanced multimodal understanding and reasoning capabilities inherent in frontier Vision-Language Models (VLMs), such as Gemini 2.0. The generalized comprehension afforded by these foundation models, with their ability to interpret visual inputs and complex text instructions, forms a powerful foundation for building embodied agents. This endeavor hinges on two fundamental components. First, Gemini needs to acquire robust embodied reasoning, gaining the ability to understand the rich geometric and temporal-spatial details of the physical world. Second, we must ground this embodied reasoning in the physical world by enabling Gemini to speak the language of physical actions, understanding contact physics, dynamics, and the intricacies of real-world interactions. Ultimately, these pieces must coalesce to enable fast, safe and dexterous control of robots in the real world.
为此,我们推出了 Gemini Robotics 系列具身 AI 模型,该系列基于我们最先进的多模态基础模型 Gemini 2.0 构建。我们首先通过一个新的开源通用具身推理基准 ERQA,验证了基础 Gemini 2.0 固有的具身推理能力的性能和通用性。然后我们介绍了两个模型:第一个模型是 Gemini Robotics-ER,这是一个以强大具身推理能力为核心的 VLM,在广泛的具身推理任务中展现出泛化能力,同时保持其核心基础模型能力。Gemini Robotics-ER 在理解物理世界至关重要的多种能力上表现出色,从 3D 感知到详细指向,再到通过代码进行机器人状态估计和可供性预测。第二个模型是 Gemini Robotics,这是一个最先进的视觉-语言-动作(VLA)模型,将强大的具身推理先验与真实世界机器人的灵巧低级控制连接起来,以解决具有挑战性的操作任务。作为通用 VLA,Gemini Robotics 可以执行各种多样且复杂的任务,同时紧密遵循语言指导,并泛化到指令、视觉和动作的分布变化。为了强调 Gemini Robotics 模型的灵活性和通用性,我们还引入了一个可选的专门化阶段,展示了 Gemini Robotics 如何适应极端灵巧性、在困难泛化设置中进行高级推理以及控制全新机器人实体。最后,我们讨论了训练大型机器人模型(如 Gemini Robotics 模型)的安全影响,并为如何在 VLA 背景下研究此类挑战提供了指导方针。具体而言,本报告重点介绍:
To this end, we introduce the Gemini Robotics family of embodied AI models, built on top of Gemini 2.0, our most advanced multimodal foundation model. We first validate the performance and generality of the base Gemini 2.0’s innate embodied reasoning capabilities with a new open-source general embodied reasoning benchmark, ERQA. We then introduce two models: The first model is Gemini Robotics-ER, a VLM with strong embodied reasoning capabilities at its core, exhibiting generalization across a wide range of embodied reasoning tasks while also maintaining its core foundation model capabilities. Gemini Robotics-ER exhibits strong performance on multiple capabilities critical for understanding the physical world, ranging from 3D perception to detailed pointing to robot state estimation and affordance prediction via code. The second model is Gemini Robotics, a state-of-the-art Vision-Language-Action (VLA) model that connects strong embodied reasoning priors to dexterous low-level control of real-world robots to solve challenging manipulation tasks. As a generalist VLA, Gemini Robotics can perform a wide array of diverse and complicated tasks, while also closely following language guidance and generalizing to distribution shifts in instructions, visuals, and motions. To emphasize the flexibility and generality of the Gemini Robotics models, we also introduce an optional specialization stage, which demonstrates how Gemini Robotics can be adapted for extreme dexterity, for advanced reasoning in difficult generalization settings, and for controlling completely new robot embodiments. Finally, we discuss the safety implications of training large robotics models such as the Gemini Robotics models, and provide guidelines for how to study such challenges in the context of VLAs. Specifically, this report highlights:
ERQA:一个专门设计用于评估多模态模型具身推理能力的开源基准,解决了缺乏超越原子能力评估的基准的问题,并促进了标准化评估和未来研究。
ERQA: An open-source benchmark specifically designed to evaluate embodied reasoning capabilities of multimodal models, addressing the lack of benchmarks that go beyond assessing atomic capabilities and facilitating standardized assessment and future research.
Gemini Robotics-ER:一个展示增强具身推理能力的视觉语言模型(VLM)。
Gemini Robotics-ER: A VLM demonstrating enhanced embodied reasoning capabilities.
Gemini Robotics:一个由机器人动作数据整合而成的视觉-语言-动作(VLA)模型,能够实现高频灵巧控制、强大的泛化能力以及跨多种机器人任务和形态的快速适应。
Gemini Robotics: A VLA model resulting from the integration of robot action data, enabling high-frequency dexterous control, robust generalization and fast adaptation across diverse robotic tasks and embodiments.
负责任开发:我们讨论并践行我们模型系列的负责任开发,遵循谷歌 AI 原则,仔细研究我们模型的社会效益和风险,以及潜在的风险缓解措施。
Responsible Development: We discuss and exercise responsible development of our family of models in alignment with Google AI Principles, carefully studying the societal benefits and risks of our models, and potential risk mitigation.
Gemini Robotics 模型是迈向更通用机器人的第一步。我们相信,最终,利用互联网规模数据中的具身推理能力,并结合来自真实世界交互的动作数据,可以使机器人深刻理解物理世界并胜任行动。这种理解将赋予它们以通用性和复杂性完成最具挑战性目标的能力,而这在以往对于机器人系统而言似乎是遥不可及的。
The Gemini Robotics models serve as an initial step towards more generally capable robots. We believe that, ultimately, harnessing the embodied reasoning capabilities from internet-scale data, grounded with action data from real-world interactions, can enable robots to deeply understand the physical world and act competently. This understanding will empower them to achieve even the most challenging goals with generality and sophistication that has so far seemed out of reach for robotic systems.
Gemini 2.0 是一种视觉-语言模型(VLM),能够超越仅需视觉理解和语言处理的任务。特别是,该模型展现出先进的具身推理(ER)能力。我们将 ER 定义为视觉-语言模型将物体和空间概念锚定在现实世界中,并综合这些信号用于下游机器人应用的能力。图 2 展示了此类能力的一些示例。在第 2.1 节中,我们首先介绍一个用于评估广泛 ER 能力的基准,并表明 Gemini 2.0 模型达到了最先进水平。在第 2.2 节中,我们展示了 Gemini 2.0 所支持的广泛具体 ER 能力。最后,在第 2.3 节中,我们展示了如何将这些能力应用于机器人应用,而无需对机器人动作数据进行微调,从而支持通过代码生成实现零样本控制以及通过上下文学习实现少样本机器人控制等用例。
Gemini 2.0 is a Vision-Language Model (VLM) that is capable of going beyond tasks that only require visual understanding and language processing. In particular, this model exhibits advanced embodied reasoning (ER) capabilities. We define ER as the ability of a Vision-Language Model to ground objects and spatial concepts in the real world, and the ability to synthesize those signals for downstream robotics applications. See some examples of such capabilities in Fig. 2. In Section 2.1, we first introduce a benchmark for evaluating a broad spectrum of ER capabilities and show that Gemini 2.0 models are state-of-the-art. In Section 2.2, we demonstrate the wide range of specific ER capabilities enabled by Gemini 2.0. Finally, in Section 2.3, we showcase how these capabilities can be put to use in robotics applications without the need for fine-tuning on robot action data, enabling use cases such as zero-shot control via code generation and few-shot robot control via in-context learning.
为了捕捉视觉语言模型(VLM)在具身推理方面的进展,我们引入了 ERQA(Embodied Reasoning Question Answering,具身推理问答),这是一个专门关注具身智能体与物理世界交互时可能所需能力的基准。ERQA 包含 400 道多项选择视觉问答(VQA)风格的问题,涵盖多种类别,包括空间推理、轨迹推理、动作推理、状态估计、指向、多视角推理和任务推理。问题类型的分布如图 4 所示。在 400 个问题中,28%的问题在提示中包含多张图像——这些需要跨图像对应概念的问题往往比单图像问题更具挑战性。
To capture progress in embodied reasoning for VLMs, we introduce ERQA, short for Embodied Reasoning Question Answering, a benchmark that focuses specifically on capabilities likely required by an embodied agent interacting with the physical world. ERQA consists of 400 multiple choice Visual Question Answering (VQA)-style questions across a wide variety of categories, including spatial reasoning, trajectory reasoning, action reasoning, state estimation, pointing, multi-view reasoning, and task reasoning. A breakdown of the distribution of question types is in Fig. 4. Of the 400 questions 28% have more than one image in the prompt — these questions that require corresponding concepts across multiple images tend to be more challenging than single-image questions.
ERQA 与现有的 VLM 基准互补,后者往往强调更原子化的能力(如物体识别、计数、定位),但在大多数情况下未充分考虑在物理世界中行动所需的更广泛能力。图 3 展示了我们 ERQA 的一些示例问题和答案。有些问题要求 VLM 跨帧识别和配准物体;另一些则要求推理物体的可供性及其与场景其余部分的 3D 关系。基准的完整细节可在 https://github.com/embodiedreasoning/ERQA 找到。
ERQA is complementary to existing VLM benchmarks, which tend to highlight more atomic capabilities (e.g., object recognition, counting, localization), but in most cases do not take sufficient account of the broader set of capabilities needed to act in the physical world. Fig. 3 shows some example questions and answers of our ERQA. Some questions require the VLM to recognize and register objects across multiple frames; others require reasoning about objects’ affordances and 3D relationships with the rest of the scene. Full details of the benchmark can be found at https://github.com/embodiedreasoning/ERQA.
我们手动标注了 ERQA 中的所有问题以确保正确性和质量。基准中的图像(而非问题)要么是我们自己拍摄的,要么来自以下数据集:OXE、UMI Data、MECCANO、HoloAssist 和 EGTEA Gaze+。在表 1 中,我们报告了 Gemini 模型和其他模型在 ERQA 上的结果,以及在 RealworldQA 和 BLINK 上的结果,这两个流行的基准也衡量空间和图像理解能力。具体来说,我们报告了 Gemini 2.0 Flash(一个强大的低延迟主力模型)和 Gemini 2.0 Pro Experimental 02-05(下文简称 Gemini 2.0 Pro Experimental,是 Gemini 在复杂任务上最好的模型)的结果。Gemini 2.0 Flash 和 Pro Experimental 在各自模型类别中在这三个基准上都取得了新的最先进成果。我们还注意到,ERQA 是这三个基准中最具挑战性的,因此这里的性能尤其值得注意。
We manually labeled all questions in ERQA to ensure correctness and quality. Images (not questions) in the benchmark are either taken by ourselves or sourced from these datasets: OXE, UMI Data, MECCANO, HoloAssist, and EGTEA Gaze+. In Table 1, we report results of Gemini models and other models on ERQA, as well as on RealworldQA and BLINK, two popular benchmarks that also measure spatial and image understanding capabilities. Specifically, we report results of Gemini 2.0 Flash, a powerful low-latency workhorse model and Gemini 2.0 Pro Experimental 02-05 (short as Gemini 2.0 Pro Experimental in the rest of the paper), the best Gemini model for complex tasks. Gemini 2.0 Flash and Pro Experimental achieve a new state-of-the-art on all three benchmarks in their respective model classes. We also note that ERQA is the most challenging benchmark across these three, making the performance here especially notable.
Gemini 2.0 模型具备高级推理能力——我们发现,如果使用思维链(CoT)提示,可以显著提高 Gemini 2.0 在基准上的性能。思维链提示鼓励模型在选择题答案之前输出推理轨迹来“思考”问题,而不是直接预测答案。我们在每个问题末尾附加以下指令作为 CoT 提示:“逐步推理答案,并展示每一步的工作。之后才给出最终答案。”结果如表 2 所示。使用 CoT 提示后,Gemini 2.0 Flash 的性能超过了未使用 CoT 的 Gemini 2.0 Pro Experimental,而 CoT 进一步提高了 Gemini 2.0 Pro Experimental 的性能。我们在图 5 中重点展示了两个这样的推理轨迹,这些问题是 Gemini 2.0 Pro Experimental 在没有 CoT 时回答错误,但在使用 CoT 后回答正确的。推理轨迹表明 Gemini 2.0 能够 1)将其空间理解精确地锚定在图像中的观察上,2)利用这种锚定来执行复杂的、逐步的具身推理。
Gemini 2.0 models are capable of advanced reasoning — we found we can significantly improve Gemini 2.0’s performance on the benchmark if we use Chain-of-Thought (CoT) prompting, which encourages the model to output reasoning traces to “think” about a problem before choosing the multiple choice answer, instead of directly predicting the answer. We use the following instruction as the CoT prompt appended at the end of each question: “Reason step by step about the answer, and show your work, for each step. Only after that, proceed to the final answer.” Results are shown in Table 2. With CoT prompting, Gemini 2.0 Flash’s performance exceeds that of Gemini 2.0 Pro Experimental without CoT, and CoT further improves Gemini 2.0 Pro Experimental’s performance. We highlight two such reasoning traces in Fig. 5, questions that Gemini 2.0 Pro Experimental answered incorrectly without CoT, but correctly with CoT. The reasoning traces demonstrate Gemini 2.0 is able to 1) precisely ground its spatial understanding in observations in the image and 2) leverage such grounding to perform complex, step-by-step embodied reasoning.
在本节中,我们将更详细地阐述 Gemini 2.0 的一些具身推理能力。我们还介绍了 Gemini Robotics-ER,这是 Gemini 2.0 Flash 的一个版本,具有增强的具身推理能力。这些能力可用于机器人应用,无需任何额外的机器人特定数据或训练。Gemini 2.0 能够理解图像中的多种 2D 空间概念。
In this section, we illustrate some of Gemini 2.0's embodied reasoning capabilities in more detail. We also introduce Gemini Robotics-ER, a version of Gemini 2.0 Flash that has enhanced embodied reasoning. These can be used in robotics applications without the need for any additional robot-specific data or training. Gemini 2.0 can understand a variety of 2D spatial concepts in images.
目标检测:Gemini 2.0 能够执行开放世界的 2D 目标检测,提供精确的 2D 边界框,其查询可以是显式的(例如,描述对象名称)或隐式的(类别、属性或功能)。
Object Detection: Gemini 2.0 can perform open-world 2D object detection, providing precise 2D bounding boxes with queries that can be explicit (e.g., describing an object name) or implicit (categories, attributes, or functions).
指向:给定任何自然语言描述,模型能够指向显式实体(如对象和对象部件),以及隐式概念(如可供性(在哪里抓取、在哪里放置)、自由空间和空间概念)。定量评估见表 3。
Pointing: Given any natural language description, the model is able to point to explicit entities like objects and object parts, as well as implicit notions such as affordances (where to grasp, where to place), free space and spatial concepts. See Table 3 for quantitative evaluations.
轨迹预测:Gemini 2.0 可以利用其指向能力生成基于其观察的 2D 运动轨迹。例如,轨迹可以基于物理运动或交互的描述。
Trajectory Prediction: Gemini 2.0 can leverage its pointing capabilities to produce 2D motion trajectories that are grounded in its observations. Trajectories can be based, for instance, on a description of the physical motion or interaction.
抓取预测:这是 Gemini Robotics-ER 中引入的新功能。它扩展了 Gemini 2.0 的指向能力,以预测自上而下的抓取。
Grasp Prediction: This is a new feature introduced in Gemini Robotics-ER. It extends Gemini 2.0's pointing capabilities to predict top-down grasps.
Gemini 2.0 还具备 3D 空间推理能力。凭借“以 3D 方式看”的能力,Gemini 2.0 能更好地理解大小、距离和方向等概念,并能利用这种理解来推理场景状态和要执行的动作。
Gemini 2.0 is also capable of 3D spatial reasoning. With the ability to “see in 3D”, Gemini 2.0 can better understand concepts like sizes, distances, and orientations, and it can leverage such understanding to reason about the state of the scene and actions to perform.
多视图对应:用图像表示 3D 信息的一种自然方式是通过多视图(例如立体)图像。Gemini 2.0 能从多视图图像中理解 3D 场景,并预测同一场景多个相机视角之间的 2D 点对应关系。
Multi-View Correspondence: A natural way of representing 3D information with images is through multi-view (e.g., stereo) images. Gemini 2.0 can understand 3D scenes from multi-view images and predict 2D point correspondences across multiple camera views of the same scene.
3D 边界框检测:这种 3D 理解也适用于单张图像——Gemini 2.0 可以直接从单目图像预测度量 3D 边界框。与 2D 检测和指向能力类似,Gemini 2.0 可以通过开放词汇描述来检测物体。
3D Bounding Box Detection: This 3D understanding applies to single images as well - Gemini 2.0 can directly predict metric 3D bounding boxes from monocular images. Like 2D Detection and Pointing capabilities, Gemini 2.0 can detect objects by open-vocabulary descriptions.
虽然可以为这些任务分别创建专家模型,但将它们融合到单个基础模型(如 Gemini 2.0)中,使得模型能够通过开放世界的自然语言指令执行具身推理任务,响应反馈并维持多轮交互。特别是,Gemini 2.0 可以将场景理解与推理相结合,以解决更复杂的任务,例如编写机器人代码(见第 2.3 节)。
While it is possible to create expert models for each of these tasks individually, fusing them in a single foundation model, such as Gemini 2.0, allows the model to perform embodied reasoning tasks with open-world natural language instructions, respond to feedback and sustain multi-turn interactions. In particular, Gemini 2.0 can combine scene understanding with reasoning to solve more complex tasks, such as writing robot code (see Section 2.3).
下面我们展示使用 Gemini 2.0 模型(Flash 和 Pro Experimental)对这些能力的详细定量和定性评估,并在适当情况下与其他 VLM 进行比较。对于某些能力,我们还提供了在 Gemini Robotics-ER 上的结果。你可以在此处找到如何提示 Gemini 2.0 以激发这些能力的代码和提示示例。
Below we present detailed quantitative and qualitative evaluations of these capabilities with Gemini 2.0 models (Flash, and Pro Experimental), as well as comparisons with other VLMs where appropriate. For some capabilities, we also present results on Gemini Robotics-ER. You can find code and prompt examples on how to prompt Gemini 2.0 to elicit these capabilities here.
**目标检测**。Gemini 2.0 能够根据自然语言查询预测二维目标边界框。在图 6 中,我们展示了使用 Gemini 2.0 Flash 在机器人可能看到的图像上进行多个二维检测的示例。Gemini 2.0 使用约定 \([y_{0},x_{0},y_{1},x_{1}]\) 来表示二维边界框。我们可以提示 Gemini 2.0 检测场景中的所有物体(图 2 中的示例)。该模型还可以通过描述来检测特定物体——例如,图 6 中的“检测所有厨房用具”。这些描述还可以包含空间线索——中间示例中的“检测图像右侧的螺母”。最后,我们可以提示 Gemini 2.0 根据物体的可供性(affordances)进行检测。在图 6 的右侧示例中,我们要求 Gemini 2.0 检测溢出物以及“可以用来清理它的东西”。Gemini 2.0 能够同时检测到溢出物和毛巾,而无需明确指定。这些示例展示了将精确的定位能力与通用视觉语言模型(VLM)相结合的优势,其中 Gemini 的开放词汇和开放世界推理实现了语义泛化的水平,这是专用专家模型难以达到的。
Object Detection. Gemini 2.0 can predict 2D object bounding boxes from natural language queries. In Fig. 6, we show multiple 2D detection examples with Gemini 2.0 Flash on images that a robot might see. Gemini 2.0 represents 2D bounding boxes with the convention \([y_{0},x_{0},y_{1},x_{1}]\). We can prompt Gemini 2.0 to detect everything in a scene (examples in Fig. 2). The model can also detect specific objects by their descriptions — for example, “detect all the kitchenware” in Fig. 6. These descriptions can contain spatial cues as well — “detecting nuts on the right side of the image” in the middle example. Finally, we can prompt Gemini 2.0 to detect objects by their affordances. In the right example of Fig. 6, we ask Gemini 2.0 to detect the spill and “what can be used to clean it up”. Gemini 2.0 is able to detect both the spill and the towel, without being specified explicitly. These examples showcase the benefit of combining precise localization capabilities with general-purpose VLMs, where Gemini’s open-vocabulary and open-world reasoning enables a level of semantic generalization that is difficult to achieve with special-purpose expert models.
**二维指向**。对于某些用例,点可以比边界框提供更灵活、更精确的图像理解和机器人控制表示。我们在各种机器人操作场景中展示了 Gemini 2.0 的指向能力(图 7)。该模型将点表示为 \([y,x]\) 元组。与二维目标检测类似,Gemini 2.0 可以指向开放词汇语言描述的任何物体。Gemini 2.0 不仅能定位整个物体,还能定位物体部件,例如勺子手柄(图 7 左)。此外,Gemini 2.0 可以指向空间概念,例如“平底锅左侧桌子上的空区域”(图 7 左)或“按照现有八个罐子的图案放置新罐子的位置”(图 7 中)。它还可以推断可供性;例如,当被要求“指向人类拿起它时会抓握的位置”时,模型正确识别了杯子手柄(图 7 右)。
2D Pointing. For some use cases, points can offer a more flexible and precise representation for image understanding and robot control than bounding boxes. We illustrate Gemini 2.0’s pointing capabilities in various robot manipulation scenes (Fig. 7). The model represents points as \([y,x]\) tuples. Similar to 2D object detection, Gemini 2.0 can point to any object described by open-vocabulary language. Gemini 2.0 can localize not only entire objects, but also object parts, such as a spoon handle (Fig. 7, left). Additionally, Gemini 2.0 can point to spatial concepts, e.g., an “empty area on the table left of the pan” (Fig. 7, left) or “where a new can should be placed following the pattern of the existing eight cans” (Fig. 7, middle). It can also infer affordances; for example, when asked to “point to where a human would grasp this to pick it up”, the model correctly identifies the mug handle (Fig. 7, right).
我们使用三个基准对 Gemini 2.0 的指向性能进行了定量评估(表 3):Paco-LVIS 用于自然图像上的物体部件指向,Pixmo-Point 用于网络图像上的开放词汇指向,Where2place 用于室内场景中的自由空间指向。关于我们如何与其他模型进行指向基准测试的详细信息,请参见第 B.2 节。Gemini 2.0 显著优于 GPT 和 Claude 等最先进的视觉语言模型(VLM)。Gemini Robotics-ER 在三个子任务中的两个上超过了专门的指向 VLM Molmo。
We quantitatively evaluate Gemini 2.0’s pointing performance in Table 3 using three benchmarks: Paco-LVIS for object part pointing on natural images, Pixmo-Point for open-vocabulary pointing on web images, and Where2place for free-space pointing in indoor scenes. See Section B.2 for details on how we benchmark pointing against other models. Gemini 2.0 significantly outperforms state-of-the-art vision-language models (VLMs) like GPT and Claude. Gemini Robotics-ER surpasses Molmo, a specialized pointing VLM, in two of the three subtasks.
**二维轨迹**。Gemini 2.0 可以利用其指向能力来预测连接多个点的二维轨迹。虽然 Gemini 2.0 无法执行复杂的运动规划(例如避开障碍物),但它仍然可以生成基于观察图像的有用轨迹。我们在图 8 中展示了一些示例。在左侧和中间的图像中,Gemini 2.0 从第一人称视频中的人手插值出一条合理的轨迹,指向它可能抓取的工具。在右侧图像中,Gemini 2.0 预测了一系列路径点,如果机器人夹爪遵循这些路径点,将擦拭托盘上的溢出区域。Gemini 2.0 的轨迹预测能力展示了关于运动和动力学的世界知识,这是机器人技术的基本能力。我们在第 4.2 节中利用这些初步的轨迹理解能力,将动作与视觉和语言能力更紧密地联系起来。
2D Trajectories. Gemini 2.0 can leverage its pointing capabilities to predict 2D trajectories that connect multiple points together. While Gemini 2.0 cannot perform complex motion planning (e.g., to avoid obstacles), it can still generate useful trajectories that are grounded in the observed images. We showcase some examples in Fig. 8. In the left and middle images, Gemini 2.0 interpolates a reasonable trajectory from a human hand in the ego-centric video to a tool that it may grasp. In the right image, Gemini 2.0 predicts a series of waypoints that, if followed by the robot gripper, would wipe the spilled area of a tray. Gemini 2.0’s trajectory prediction capabilities exhibit world knowledge about motion and dynamics which is a fundamental capability for robotics. We capitalize on these nascent trajectory understanding capabilities to tie actions to vision and language capabilities in a much stronger fashion in Section 4.2.
**自上而下的抓取**。Gemini 2.0 的语义指向能力可以自然地扩展到自上而下的抓取姿态,表示为 \(y\)、\(x\) 和旋转角 \(\theta\)。Gemini Robotics-ER 进一步改进了这一能力,如图 9 所示。例如,我们可以提示在香蕉茎上或香蕉中心进行抓取(右图)。我们在第 2.3 节中展示了如何将这种抓取预测直接用于真实机器人的下游控制。
Top-Down Grasps. Gemini 2.0’s semantic pointing capabilities can be naturally extended to top-down grasping poses, represented as \(y\), \(x\), and a rotation angle \(\theta\). This capability is further improved in Gemini Robotics-ER, as shown in Fig. 9. For example, we can prompt for a grasp either on the stem of the banana or the center of the banana (right image). We show how such grasp predictions can be directly used for downstream robot control on real robots in Section 2.3.
多视图对应。Gemini 还能理解世界的三维结构。一个例子是它从多个视图理解三维场景的能力。例如,给定一张标注了若干点的初始图像,以及同一场景从不同视角拍摄的新图像,我们可以询问 Gemini 2.0 初始图像中的哪些点在第二张图像中仍然可见,并查询这些点的坐标。从图 10 的示例中,我们观察到 Gemini 2.0 能够在差异极大的视图之间执行多视图对应。在顶部图像对中,模型正确预测红点指的是这些第一人称视角图像中人手持的物体,尽管场景其余部分的视角已发生显著变化。在底部图像对中,模型正确预测橙点在第二张图像中不可见。这种多视图理解对机器人领域很有用,机器人可以利用 Gemini 2.0 推理多个图像流(例如立体视图、头部和手腕视图),以更好地理解其观测的三维空间关系。
Multi-view Correspondence. Gemini can also understand the 3D structure of the world. One example is its ability to understand a 3D scene from multiple views. For instance, with an initial image annotated with a list of points and a new image of the same scene from a different view, we can ask Gemini 2.0 which of the points from the initial image are still visible in the second image and we can query the coordinates of those points. From the examples in Fig. 10, we observe that Gemini 2.0 can perform multi-view correspondence across dramatically different views. In the top image pair, the model correctly predicts that the red point refers to an object held by the human in these egocentric images, even though the view of the rest of the scene has changed significantly. In the bottom image pair, the model correctly predicts that the orange point is not visible in the second image. Such multi-view understanding is useful for robotics domains where a robot can use Gemini 2.0 to reason about multiple image streams (e.g., stereo views, head and wrist views) to better understand the 3D spatial relationships of its observations.
3D 检测。Gemini 2.0 还能从单张图像预测度量三维边界框。与其 2D 检测能力类似,Gemini 2.0 的 3D 检测能力也是开放词汇的,如图 11 所示。在表 4 中,我们报告了 Gemini 2.0 在 SUN-RGBD(一个用于 3D 物体检测和场景理解的流行数据集和基准)上的 3D 检测性能,并与基线专家模型(ImVoxelNet、Implicit3D 和 Total3DUnderstanding)进行比较。Gemini 2.0 的 3D 检测性能与现有最先进的专家模型相当,其中 Gemini Robotics-ER 在 SUN-RGBD 基准上取得了新的最先进结果。虽然这些基线模型处理的是封闭类别集,但 Gemini 支持开放词汇查询。
3D Detection. Gemini 2.0 can also predict metric 3D bounding boxes from single images. Similar to its 2D detection capabilities, Gemini 2.0’s 3D detection capability is also open-vocabulary, as illustrated in Fig. 11. In Table 4, we report Gemini 2.0’s 3D detection performance using SUN-RGBD, a popular dataset and benchmark for 3D object detection and scene understanding, and compare it with baseline expert models (ImVoxelNet, Implicit3D, and Total3DUnderstanding). Gemini 2.0’s 3D detection performance is comparable to existing state-of-the-art expert models, with Gemini Robotics-ER achieving a new state-of-the-art on the SUN-RGBD benchmark. While these baselines work with a closed set of categories, Gemini allows for open-vocabulary queries.
Gemini 2.0 的具身推理能力使得无需任何机器人动作数据训练即可控制机器人。它可以开箱即用地执行所有必要步骤:感知、状态估计、空间推理、规划和控制。以往的工作需要组合多个模型才能实现这一目标,而 Gemini 2.0 将所需能力统一于单一模型之中。
Gemini 2.0’s embodied reasoning capabilities make it possible to control a robot without it ever having been trained with any robot action data. It can perform all the necessary steps, perception, state estimation, spatial reasoning, planning and control, out of the box. Whereas previous work needed to compose multiple models to this end, Gemini 2.0 unites all required capabilities in a single model.
下面我们研究两种不同的方法:通过代码生成的零样本控制,以及通过上下文学习(下文简称“ICL”)的少样本控制——即我们以少量上下文演示来条件化模型,使其学习新行为。Gemini Robotics-ER 在两种设置下的一系列不同任务上都取得了良好性能,我们发现尤其是零样本机器人控制性能与更好的具身理解密切相关:Gemini Robotics-ER 为此接受了更全面的训练,其任务完成率相比 Gemini 2.0 提升了近 2 倍。
Below we study two distinct approaches: zero-shot robot control via code generation, and few-shot control via in-context learning (also denoted as “ICL” below) - where we condition the model on a handful of in-context demonstrations for a new behavior. Gemini Robotics-ER achieves good performance across a range of different tasks in both settings, and we find that especially zero-shot robot control performance is strongly correlated with better embodied understanding: Gemini Robotics-ER, which has received more comprehensive training to this end, improves task completion by almost 2x compared to Gemini 2.0.
通过代码生成的零样本控制。为测试 Gemini 2.0 的零样本控制能力,我们将其固有的代码生成能力与第 2.2 节所述的具身推理能力相结合。我们在双臂 ALOHA 2 机器人上进行了实验。为控制机器人,Gemini 2.0 可访问一个 API,该 API 能够将每个夹爪移动到指定姿态、打开和关闭每个夹爪,并提供当前机器人状态的读数。该 API 还提供感知功能;无需调用外部模型,而是由 Gemini 2.0 自身检测物体边界框、物体上的点,并生成如第 2.2 节所述的俯视抓取姿态。
Zero-shot Control via Code Generation. To test Gemini 2.0’s zero-shot control capabilities, we combine its innate ability to generate code with the embodied reasoning capabilities described in Section 2.2. We conduct experiments on a bimanual ALOHA 2 robot. To control the robot, Gemini 2.0 has access to an API that can move each gripper to a specified pose, open and close each gripper, and provide a readout of the current robot state. The API also provides functions for perception; no external models are called, instead Gemini 2.0 itself detects object bounding boxes, points on objects, and generates the top down grasp pose as described in Section 2.2.
在一个回合中,Gemini 2.0 首先接收系统提示、机器人 API 描述和任务指令。然后,Gemini 2.0 迭代地接收显示当前场景状态、机器人状态和执行反馈的图像,并输出在环境中执行以控制机器人的代码。生成的代码使用 API 来理解场景并移动机器人,执行循环使 Gemini 2.0 能够在必要时做出反应并重新规划(例如,图 34)。API 和回合控制流程的概览见图 12。
During an episode, Gemini 2.0 is initially passed a system prompt, a description of the robot API, and the task instructions. Then Gemini 2.0 iteratively takes in images that show the current state of the scene, the robot state, and execution feedback, and outputs code that is executed in the environment to control the robot. The generated code uses the API to understand the scene and move the robot and the execution loop allows Gemini 2.0 to react and replan when necessary (e.g., Fig. 34). An overview of the API and episodic control flow is given in Fig. 12.
通过上下文示例的少样本控制。前述结果展示了 Gemini Robotics-ER 如何被有效用于完全零样本地处理一系列任务。然而,一些灵巧操作任务超出了 Gemini 2.0 当前零样本执行的能力。受此类情况启发,我们证明了该模型可以通过少量上下文演示进行条件化,并立即模仿这些行为。与前述示例中生成代码不同,我们转而提示模型直接生成末端执行器姿态的轨迹,遵循演示中的示例。
Few-shot control via in-context examples. The previous results demonstrated how Gemini Robotics-ER can be effectively used to tackle a series of tasks entirely zero-shot. However, some dexterous manipulation tasks are beyond Gemini 2.0’s current ability to perform zero-shot. Motivated by such cases, we demonstrate that the model can be conditioned on a handful of in-context demonstrations, and can then immediately emulate those behaviors. Instead of generating code, as in the previous examples, we instead prompt the model to generate trajectories of end-effectors poses directly, following the examples in the demonstrations.
我们扩展了文献[引用]中提出的方法,该方法将 \(k\) 条遥操作的机器人动作轨迹转换为物体和末端执行器位姿的列表,将其标记为文本并添加到提示中(图 13)。得益于 Gemini Robotics-ER 的具身推理能力,我们不需要任何外部模型来提取视觉关键点和物体位姿(如参考工作所做的那样);Gemini Robotics-ER 可以自行完成。除了观察和动作之外,我们还穿插了以语言描述的执行动作,这些描述在推理时引发模型的推理。模型模仿了上下文轨迹中的自然语言推理,并变得更好,例如,理解何时使用哪只手臂,或更准确地预测与物体交互的位置。使用大型多模态模型的一个优势是能够根据观察、动作和语言来调节其行为,而所有这些的组合优于任何单一模态。
We extend the method proposed in [reference], which translates \(k\) teleoperated trajectories of robot actions into a list of objects and end-effector poses, tokenizing them as text and adding them to the prompt (Fig. 13). Thanks to the embodied reasoning abilities of Gemini Robotics-ER, we do not need any external models to extract visual keypoints and object poses (as was done in the referenced work); Gemini Robotics-ER can do this itself. In addition to observations and actions, we interleave descriptions of the performed actions in language that elicits reasoning at inference time in the model. The model emulates the natural language reasoning from the in-context trajectories and becomes better at, for example, understanding which arm to use when, or more accurately predicting where to interact with objects. One advantage of using a large multimodal model is the ability to condition its behavior on observations, actions, and language, with the combination of all outperforming any modality in isolation.
使用这种方法(10 个演示)的结果如表 5 和表 6 所示。Gemini 2.0 Flash 和 Gemini Robotics-ER 都能有效地完全在上下文中使用演示来提高性能。Gemini 2.0 Flash 在仿真中的性能达到 51%,而 Gemini Robotics-ER 在仿真和真实世界中均达到 65%。与零样本代码生成方法相比,性能提升主要来自更灵巧的任务,如物体交接、折叠连衣裙或打包玩具,在这些任务中,演示可以调节模型输出更精确的双臂轨迹。
The results using this approach (with 10 demonstrations) are shown in Table 5 and Table 6. Both Gemini 2.0 Flash and Gemini Robotics-ER are able to effectively use demonstrations entirely in-context to improve performance. Gemini 2.0 Flash’s performance reaches 51% in simulation, and Gemini Robotics-ER achieves 65% in both simulation and the real world. Most of the performance improvements with respect to the zero-shot code generation approach come from more dexterous tasks, like handover of objects, folding a dress, or packing a toy, where demonstrations can condition the model to output more precise, bimanual trajectories.
这组实验表明,Gemini 2.0 Flash 及其 ER 增强变体 Gemini Robotics-ER 可以直接用于控制机器人,作为感知模块(例如,物体检测)、规划模块(例如,轨迹生成)和/或通过生成和执行代码来编排机器人运动。它还显示了具身推理能力的模型性能与下游机器人控制之间的强相关性。同时,我们的实验表明,该模型还能够利用上下文学习的力量,仅从少量演示中学习,并通过直接输出末端执行器位姿的轨迹来提高更灵巧和双臂任务(如折叠衣物)的性能。然而,作为 VLM,对于机器人控制存在固有的局限性,尤其是对于更灵巧的任务,因为需要中间步骤将模型固有的具身推理能力与机器人动作连接起来。在下一节中,我们将介绍 Gemini Robotics,一个端到端的视觉-语言-动作模型,它能够实现更通用和更灵巧的机器人控制。
This set of experiments suggests that Gemini 2.0 Flash and its ER enhanced variant, Gemini Robotics-ER, can be used directly to control robots, as a perception module (e.g., object detection), a planning module (e.g., trajectory generation), and/or to orchestrate robot movements by generating and executing code. It also shows strong correlation between the model performance of embodied reasoning capabilities and the downstream robotic control. At the same time, our experiments demonstrate that the model is also able to tap into the power of in-context learning to learn from just a few demonstrations and boost performance on more dexterous and bimanual tasks, such as folding clothes, by directly outputting trajectories of end-effector poses. However, as a VLM, there are inherent limitations for robot control, especially for more dexterous tasks, due to the intermediate steps needed to connect the model’s innate embodied reasoning capabilities to robotic actions. In the next section, we will introduce Gemini Robotics, an end-to-end Vision-Language-Action Model that enables more general-purpose and dexterous robot control.
在本节中,我们介绍 Gemini Robotics,它是 Gemini 的一个衍生模型,经过微调可直接预测机器人动作。Gemini Robotics 是一个通用模型,能够解决不同环境中的灵巧任务,并支持不同的机器人形态。我们首先研究模型在包含动作标注的机器人数据以及其他多模态数据的大规模多样化数据集上训练后的表现。得到的模型可以开箱即用地解决多种短时程灵巧任务(第 3.2 节),紧密遵循自然语言指令(第 3.3 节),并继承了 Gemini Robotics-ER 的泛化能力,对场景视觉变化、物体位置和实例表现出鲁棒性(第 3.4 节)。在第 4 节中,我们进一步测试 Gemini Robotics 的极限,将其专门用于具有挑战性的高灵巧长时程任务(第 4.1 节)以及更极端的泛化场景(第 4.2 节)。我们还研究了对新灵巧任务的快速适应(第 4.3 节)以及对具有全新外形、动作和观测的形态的适应(第 4.4 节)。
In this section, we present Gemini Robotics, a derivative of Gemini that has been fine-tuned to predict robot actions directly. Gemini Robotics is a general-purpose model capable of solving dexterous tasks in different environments and supporting different robot embodiments. We first study the model after training on a large and diverse dataset consisting of action-labeled robot data as well as other multimodal data. The resulting model can solve a large variety of short-horizon dexterous tasks out of the box (Section 3.2), closely follows natural language instructions (Section 3.3) and inherits Gemini Robotics-ER generalization capabilities, showing robustness to visual variations of the scene, object positions and instances (Section 3.4). In Section 4, we further test the limits of Gemini Robotics, and specialize it to challenging highly dexterous long-horizon tasks (Section 4.1), and to more extreme generalization scenarios (Section 4.2). We also investigate rapid adaptation to novel dexterous tasks (Section 4.3) as well as adaptation to embodiments with completely new form factors, actions and observations (Section 4.4).
**模型**。像 Gemini Robotics-ER 这样的大型 VLM 的推理通常较慢,并且需要特殊硬件。这在 VLA 模型的背景下可能会引发问题,因为推理可能无法在机载运行,且由此产生的延迟可能与实时机器人控制不兼容。Gemini Robotics 旨在解决这些挑战。它由两个组件组成:一个托管在云端的 VLA 主干(Gemini Robotics 主干)和一个在机器人机载计算机上运行的本地动作解码器(Gemini Robotics 解码器)。Gemini Robotics 主干由 Gemini Robotics-ER 的蒸馏版本构成,其查询到响应的延迟已从数秒优化至 160 毫秒以下。机器人上的 Gemini Robotics 解码器补偿了主干的延迟。当主干和本地解码器结合时,从原始观测到低级动作块的端到端延迟约为 250 毫秒。由于动作块中包含多个动作,有效控制频率为 50Hz。整个系统不仅能在主干延迟的情况下产生平滑的运动和反应性行为,还保留了主干的泛化能力。我们的模型架构概览见图 14。
Model. Inference in large VLMs like Gemini Robotics-ER is often slow and requires special hardware. This can cause problems in the context of VLA models, since inference may not be feasible to be run onboard, and the resulting latency may be incompatible with real-time robot control. Gemini Robotics is designed to address these challenges. It consists of two components: a VLA backbone hosted in the cloud (Gemini Robotics backbone) and a local action decoder running on the robot's onboard computer (Gemini Robotics decoder). The Gemini Robotics backbone is formed by a distilled version of Gemini Robotics-ER and its query-to-response latency has been optimized from seconds to under 160ms. The on-robot Gemini Robotics decoder compensates for the latency of the backbone. When the backbone and local decoder are combined, the end-to-end latency from raw observations to low-level action chunks is approximately 250ms. With multiple actions in the chunk, the effective control frequency is 50Hz. The overall system not only produces smooth motions and reactive behaviors despite the latency of the backbone, but also retains the backbone's generalization capabilities. An overview of our model architecture is available in Fig. 14.
**数据**。我们在 12 个月的时间里,在一组 ALOHA 2 机器人上收集了大规模遥操作机器人动作数据集,其中包含数千小时的现实世界专家机器人演示。该数据集包含数千个多样化任务,涵盖了不同的操作技能、物体、任务难度、回合长度和灵巧性要求。训练数据还包括非动作数据,如网络文档、代码、多模态内容(图像、音频、视频)以及具身推理和视觉问答数据。这提高了模型在众多机器人任务和请求中的理解、推理和泛化能力。
Data. We collected a large-scale teleoperated robot action dataset on a fleet of ALOHA 2 robots over 12 months, which consists of thousands of hours of real-world expert robot demonstrations. This dataset contains thousands of diverse tasks, covering scenarios with varied manipulation skills, objects, task difficulties, episode horizons, and dexterity requirements. The training data further includes non-action data such as web documents, code, multi-modal content (image, audio, video), and embodied reasoning and visual question answering data. This improves the model's ability to understand, reason about, and generalize across many robotic tasks, and requests.
**基线**。我们将 Gemini Robotics 与两个最先进的模型进行比较:第一个是 \(\pi_{0}\) 重实现,即我们对开放权重的最先进 \({\pi_{0}}\) VLA 模型的重实现。我们在多样化的训练混合数据上训练 \(\pi_{0}\) 重实现,并发现该模型优于作者发布的公开检查点,因此将其作为我们实验中最具性能的 VLA 基线(详见第 C.2 节)。第二个是多任务扩散策略(受 ALOHA Unleashed 启发,但修改为任务条件化),该模型已被证明在从多模态演示中学习灵巧技能方面有效。两个基线均使用我们多样化数据混合的相同组成进行训练直至收敛。Gemini Robotics 主要在云端运行,并带有本地动作解码器,而两个基线则在配备 Nvidia RTX 4090 GPU 的工作站上本地运行。本节呈现的所有实证证据均基于严格的真实世界机器人实验,包括 A/B 测试和统计分析(详见第 C.1 节)。
Baselines. We compare Gemini Robotics to two state-of-the-art models: The first one is \(\pi_{0}\) re-implement, which is our re-implementation of the open-weights state-of-the-art \({\pi_{0}}\) VLA model. We train \(\pi_{0}\) re-implement on our diverse training mixture and find this model to outperform the public checkpoint released by the authors, and hence, report it as the most performant VLA baseline in our experiments (see Section C.2 for more details). The second is a multi-task diffusion policy (inspired by ALOHA Unleashed but modified to be task-conditioned), a model that has been shown to be effective in learning dexterous skills from multi-modal demonstrations. Both baselines were trained to convergence using the same composition of our diverse data mixture. Gemini Robotics runs primarily in the cloud with a local action decoder, whereas both baselines run locally on a workstation equipped with an Nvidia RTX 4090 GPU. All empirical evidence presented in this section is based on rigorous real-world robot experiments, with A/B testing and statistical analysis (more details in Section C.1).
在我们的第一组实验中,我们展示了 Gemini Robotics 能够解决各种灵巧操作任务。我们在短视界灵巧任务上评估了该模型的性能,并与最先进的多任务基线进行了比较。我们对所有模型进行开箱即用的评估,即不进行任何任务特定的微调或额外的提示,在从第 3.1 节数据集中采样的 20 个任务上进行。我们选择了多样化的场景设置(其中一些如图 15 所示),涵盖洗衣房(例如“折叠裤子”)、厨房(例如“叠放量杯”)、杂乱的办公桌(例如“打开粉色文件夹”)以及其他日常活动(例如“打开眼镜盒”)。这些选定的任务还需要不同程度的灵巧性——从简单的抓取和放置(例如“从桌子中央拿起鞋带”)到需要双手协调的变形物体灵巧操作(例如“将电线缠绕在耳机上”)。我们在图 15 中展示了模型在这些任务上的 rollout 示例,并在第 C.1.1 节中列出了完整任务列表。
In our first set of experiments, we demonstrate that Gemini Robotics can solve a wide range of dexterous tasks. We evaluate the performance of this model on short-horizon dexterous tasks, and compare to state-of-the-art multi-task baselines. We evaluate all models out of the box, i.e., without any task-specific fine-tuning or additional prompting, on 20 tasks sampled from our dataset in Section 3.1. We choose diverse scene setups (some of them illustrated in Fig. 15), spanning a laundry room (e.g., “fold pants”), kitchen (e.g., “stack measuring cup”), cluttered office desk (e.g., “open pink folder”), and other day-to-day activities (e.g., “open glasses case”). These selected tasks also require varying levels of dexterity – from simple pick-and-place (e.g., “pick the shoe lace from the center of the table”) to dexterous manipulation of deformable objects that requires two-hand coordination (e.g., “wrap the wire around the headphone”). We show examples of our model rollouts of these tasks in Fig. 15 and full list of tasks in Section C.1.1.
图 16 总结了我们的模型和基线的性能。我们发现 Gemini Robotics 模型在开箱即用的情况下能熟练完成一半的任务,成功率超过 \(80\%\) 。值得注意的是,我们的模型在变形物体操作(“折叠粉色布料”、“将电线缠绕在耳机上”)方面表现出色,而基线在这些任务上表现不佳。对于更具挑战性的任务(例如“打开粉色文件夹”、“插入红色积木”、“将电线缠绕在耳机上”),我们发现 Gemini Robotics 是唯一能够实现非零成功率的方法,这凸显了高容量模型架构与跨所有模态(视觉、语言和动作)的高质量多样化数据的结合对于多任务策略学习至关重要。最后,我们发现一些最灵巧的任务仅从多任务设置中学习仍然相当困难(例如“插入鞋带”):我们将在第 4.1 节讨论 Gemini Robotics 解决这些任务以及更长视界挑战性任务的专业化方案。
Fig. 16 summarizes the performance of our model and the baselines. We find that the Gemini Robotics model is proficient at half of the tasks out of the box with a success rate exceeding \(80\%\) . Notably, our model excels at deformable object manipulation ( “fold pink cloth”, “wrap the wire around the headphone”), while the baselines struggle with these tasks. For the more challenging tasks, (e.g., “open pink folder”, “insert red block”, “wrap the wire around the headphone”), we find that Gemini Robotics is the only method that can achieve non-zero success, highlighting that a combination of a high-capacity model architecture along with high-quality diverse data across all modalities (vision, language, and action) is essential for multi-task policy learning. Finally, we find that some of the most dexterous tasks are still quite challenging to learn purely from the multi-task setup (e.g., “insert shoe lace”): we discuss our specialization recipe for Gemini Robotics to solve these and longer-horizon challenging tasks in Section 4.1.
第二组实验测试模型遵循自然语言指令的能力。我们选取了 25 条语言指令,在五个不同的评估场景中进行评估,包括训练场景以及包含未见物体和容器的全新场景(详见第 C.1.2 节)。评估重点在于必须精确遵循的语言命令(例如,“将蓝色夹子放在黄色便利贴的右侧”)——与“清理桌子”这类开放式抽象指令形成对比。我们在图 17 中可视化了轨迹并报告了二元任务成功率。
The second set of experiments tests the model’s ability to follow natural language instructions. We select 25 language instructions to be evaluated in five diverse evaluation scenes, including training scenes as well as novel scenes with unseen objects and receptacles (details in Section C.1.2). The evaluation focuses on language commands that must be precisely followed (e.g., “Place the blue clip to the right of the yellow sticky notes”) – in contrast to open-ended abstract instructions like “clean the table”. We visualize rollouts and report the binary task success rates in Fig. 17.
我们的实验表明,强大的可操控性源于高质量多样化数据与强大视觉语言骨干网络的结合。Gemini Robotics 和重新实现的\(\pi_{0}\)在简单的分布内场景中也优于扩散基线,这表明需要强大的语言编码器。然而,特别是在包含新物体和细粒度指令的挑战性场景中(例如,“将牙膏放入收纳盒的底部隔间”),我们发现 Gemini Robotics 比任一基线都更有效(图 17)。虽然基于 PaliGemma 的重新实现的\(\pi_{0}\)能正确接近训练中见过的物体,但它在解释描述性语言属性(例如,“顶部黑色容器”、“蓝色夹子”)方面存在困难,并且无法解决包含未见物体和语言描述符的任务。
Our experiments suggest that strong steerability arises from a combination of high-quality diverse data and a capable vision-language backbone. Gemini Robotics and the re-implemented \(\pi_{0}\) outperform the diffusion baseline, even in simple in-distribution scenes, suggesting that a strong language encoder is required. However, especially in challenging scenes with novel objects and fine-grained instructions (e.g., “Place the toothpaste in the bottom compartment of the caddy”), we find that Gemini Robotics is more effective than either baseline (Fig. 17). While the PaliGemma-based re-implemented \(\pi_{0}\) correctly approaches objects that were seen during training, it struggles with interpreting descriptive language attributes (e.g., “top black container”, “blue clip”) and fails to solve tasks with unseen objects and language descriptors.
缺乏稳健的泛化能力是机器人在家庭和工业应用中大规模部署的关键瓶颈。在最后一组实验中,我们评估了 Gemini Robotics 在处理先前工作中被认为重要的三个轴向上的变化的能力。
Lack of robust generalization is a key bottleneck for large-scale deployment of robots in domestic and industrial applications. In the final set of experiments, we evaluate Gemini Robotics’s ability to deal with variations along three axes that have been considered important in prior work.
视觉泛化:模型应对场景的视觉变化保持不变性,这些变化不影响解决任务所需的动作。这些视觉变化可以包括背景、光照条件、干扰物体或纹理的变化。
Visual Generalization: The model should be invariant to visual changes of the scene that do not affect the actions required to solve the task. These visual changes can include variations in background, lighting conditions, distractor objects, or textures.
指令泛化:模型应理解自然语言指令中的不变性和等价性。超越第 3.3 节研究的细粒度可控性,模型应理解释义、对拼写错误具有鲁棒性、理解不同语言以及不同级别的具体性。
Instruction Generalization: The model should understand invariance and equivalence in natural language instructions. Going beyond fine-grained steerability studied in Section 3.3, the model should understand paraphrasing, be robust to typos, understand different languages, and varying levels of specificities.
动作泛化:模型应能够适应已学习的动作或合成新的动作,例如泛化到训练中未见过的初始条件(如物体放置)或物体实例(如形状或物理属性)。
Action Generalization: The model should be capable of adapting learned movements or synthesizing new ones, for instance to generalize to initial conditions (e.g., object placement) or object instances (e.g., shape or physical properties) not seen during training.
我们使用多样化的任务套件评估 Gemini Robotics 和基线的泛化性能。该基准共包含 85 个任务,其中 20% 在训练分布内,28% 评估视觉泛化,28% 评估指令泛化,24% 评估动作泛化。图 18-图 20 展示了我们任务套件中三种不同类型变化的示例。详细的任务分解请参见第 C.1.3 节。图 21 报告了平均进度分数。该指标比二元任务成功提供了更连续的度量,使我们能够以更细粒度可视化每个策略的进度,尤其是困难任务的进度(每个任务的进度分数定义见附录 C.1.3.3)。我们还在附录的图 40 中提供了相同图表的成功率版本。
We evaluate the generalization performance of Gemini Robotics and the baselines using a diverse task suite. This benchmark consists of 85 tasks in total, of which 20% are within the training distribution, 28% evaluate visual generalization, 28% evaluate instruction generalization, and 24% evaluate action generalization. Fig. 18 - Fig. 20 show examples of the three different types of variations in our task suite. For a detailed breakdown of tasks, please see Section C.1.3. Fig. 21 reports average progress scores. This metric provides a more continuous measure than the binary task success, and gives us the finer granularity to visualize the policies’ progress of each task, especially the hard ones (progress score for each task is defined in Appendix C.1.3.3). We also provide the same plot in success rate in Fig. 40 in the Appendix.
如图 21 所示,Gemini Robotics 始终优于基线模型,并且更有效地处理了所有三种类型的变体。即使在基线模型完全失效的情况下(例如,使用新语言的指令),Gemini Robotics 也能取得非零性能。我们推测,这些改进得益于更大、更强大的 VLM 主干网络,包括 Gemini 2.0 中使用的先进视觉编码器,以及多样化的训练数据。
Gemini Robotics consistently outperforms the baselines and handles all three types of variations more effectively, as shown in Fig. 21. Gemini Robotics even achieves non-zero performance in cases where the baselines fail catastrophically, e.g., instructions in a new language. We speculate that these improvements result from the larger and more powerful VLM backbone, including the state-of-the-art vision encoder used in Gemini 2.0, combined with diverse training data.
Gemini Robotics 模型是一个强大的机器人通才,能够解决一系列灵巧操作任务,并展现出开箱即用的非平凡泛化能力。在本节中,我们进一步测试该模型的极限,并探索未来进一步提升其通才能力的可能途径。具体而言,我们(1)测试模型通过进一步专业化,在更具挑战性的长时程灵巧任务上变得熟练的能力;(2)通过基于语义的具身推理优化其泛化能力。我们还探索(3)快速适应新任务和新环境的可能性,以及(4)适应新机器人形态的能力。其中(1)和(2)为未来模型改进提供了重要信息,而(3)和(4)则是模型实际部署所需的理想特性。
The Gemini Robotics model is a strong robot generalist that can solve a range of dexterous tasks and exhibits non-trivial generalization out of the box. In this section, we further test the limits of the model and explore possible avenues for further improving its generalist capabilities in the future. In particular, we (1) test the model’s ability to become proficient at much more challenging long-horizon dexterous tasks with further specialization, and (2) optimize its capacity for generalization through semantically-grounded embodied reasoning. We also explore (3) the possibility of rapid adaptation to novel tasks and environments, (4) as well as the adaptation to new robot embodiments. Whereas (1,2) provide important information for future model improvements, (3) and (4) are desired properties for practical deployment of the model.
在第 3.2 节中,我们展示了 Gemini Robotics 模型可以开箱即用地完成短时程灵巧任务。在此,我们表明,使用一组狭窄的高质量数据对模型进行微调,可以使模型专门化,以解决高度灵巧、具有挑战性的长时程任务,这些任务在难度上超出了通用模型的范围。具体来说,我们选择了六个任务(图 22)来展示我们模型在专门化后的各种能力:
In Section 3.2, we showed that the Gemini Robotics model can accomplish short-horizon dexterous tasks out of the box. Here, we show that fine-tuning the model with a narrow set of high-quality data can specialize the model to solve highly dexterous, challenging, long-horizon tasks that are, in terms of their difficulty, beyond the scope of the generalist model. In particular, we select six tasks (Fig. 22) to demonstrate the various capabilities of our model after specialization:
制作折纸狐狸:机器人需要将一张纸折叠成狐狸头的形状。该任务需要 4 次精确的折叠,每次折叠都需要对齐、弯曲、捏紧和压痕,且纸层数量不断增加。这需要非常精确和可靠的双臂协调,因为即使是一个小错误也可能导致不可恢复的失败。
Make an origami fox: The robot needs to fold a paper into the shape of a fox's head. This task needs 4 precise folds, each requiring aligning, bending, pinching, and creasing, with an increasing number of paper layers. This requires very precise and reliable bi-arm coordination, as even a small error can lead to an irrecoverable failure.
打包午餐盒:机器人需要将几件物品装入午餐袋:首先,它需要将一片面包插入塑料袋的窄缝中,拉上拉链,然后将这个塑料袋和一根能量棒转移到午餐袋中。接下来,它必须将葡萄转移到一个容器中,盖上盖子,并将容器移入午餐袋。最后,机器人必须拉上午餐袋的拉链。其中几个子任务(例如,插入面包、关闭容器盖、拉上午餐袋拉链)需要双臂之间的精确协调和精细的夹爪运动。
Pack a lunch-box: The robot needs to pack a lunch bag with several items: It first needs to insert a slice of bread into the narrow slit of a plastic bag, zip it, and transfer this plastic bag and an energy bar into the lunch bag. Next, it must transfer the grapes into a container, seal its lid, and move the container into the lunch bag. Finally, the robot must zip the lunch bag close. Several of the subtasks (e.g., inserting the bread, closing the container lid, zipping the lunch bag) require precise coordination between the two arms and fine gripper motion.
拼字棋盘游戏:在这个游戏中,人类在机器人面前放置(或绘制)一个物体的图片。机器人必须识别该物体,并通过将字母块移动到棋盘上来物理拼出描述该物体的三个字母单词。该任务需要视觉识别以及紧密的视觉-语言-动作对齐。
Spelling board game: In this game, the human places (or draws) a picture of an object in front of the robot. The robot must identify the object and physically spell a three-letter word describing the object by moving alphabet tiles onto a board. This task requires visual recognition, and tight vision-language-action grounding.
玩纸牌游戏:机器人必须使用自动发牌机抽取三张牌,并将它们转移到另一只手中。然后机器人必须等待人类出牌,然后从手中打出一张牌,最后弃牌。这是一个具有挑战性的精细操作任务,要求机器人交接薄薄的扑克牌,并精确地从手中挑选一张牌。
Play a game of cards: The robot must use an automatic card dealer machine to draw three cards and transfer them to its other hand. The robot must then wait for the human to play, then play a card from its hand, and finally, fold its hand. This is a challenging fine-grained manipulation task that requires the robot to handover thin playing cards and precisely pick a card from its hand.
将荷兰豆加入沙拉:机器人必须使用金属夹子从碗中夹起荷兰豆,并将其放入另一个碗中。使用夹子需要双臂协调:一只手臂握住夹子,另一只手臂施加压力以抓取和释放豌豆。
Add snap peas to salad: The robot must use metal tongs to grab snap peas from a bowl and add them to a different bowl. Using tongs requires bi-manual coordination: One arm holds the tongs while the other one applies pressure to grasp and release the peas.
将坚果加入沙拉:机器人必须使用勺子将坚果从垂直容器中舀出并倒入沙拉碗。舀取动作需要灵巧性,才能成功地从较高的容器中收集坚果,然后将其倒入沙拉碗中。
Add nuts to salad: The robot must use a spoon to scoop nuts from a vertical container to the salad bowl. The scooping motion requires dexterity to successfully collect nuts from the taller container and then pour them into the salad bowl.
我们为每个任务整理了 2000 到 5000 条高质量示范数据,并使用每个专门化数据集对第 3 节中的 Gemini Robotics 检查点进行微调。我们将这些专门化模型的性能与基线模型的专门化版本(\(\pi_{0}\) 重实现专门化模型和多任务扩散专门化模型)进行比较,两者都在相同的数据集上进行了微调。此外,为了评估第 3 节中使用的多样化训练数据的重要性,我们从头训练了一个单任务扩散策略和另一个 Gemini Robotics 专门化模型,而不是使用第 3 节中的检查点。我们在真实世界中广泛评估了所有模型,并在图 23 中报告了任务成功率(进度分数结果见附录图 42)。除拼写棋盘游戏外,每个任务对每个模型进行 20 次试验,拼写棋盘游戏进行了 12 次试验。
We curate between 2000 and 5000 episodes of high-quality demonstration data for each task, and fine-tune the Gemini Robotics checkpoint from Section 3 using each specialization dataset. We compare the performance of these specialist models with specialized versions of the baselines ( \(\pi_{0}\) re-implement specialist and Multi-task diffusion specialist), both of which are fine-tuned on the same datasets. Additionally, to evaluate the importance of diverse training data used in Section 3, we train a single task diffusion policy and another Gemini Robotics specialist from scratch instead of from the checkpoints from Section 3. We evaluate all models extensively in the real-world and report task success rate in Fig. 23 (progress score results available in Appendix in Fig. 42). We conduct 20 trials per task for each model for all tasks except for the spelling board game, for which 12 trials are conducted.
我们发现,我们的专门化模型能够以 79%的平均成功率解决所有这些任务。最值得注意的是,它在完整的长时间午餐盒打包任务上达到了 100%的成功率,该任务耗时超过 2 分钟。在拼写游戏中,它能够正确读取和拼写印刷图像中的单词(这些图像出现在专门化数据集中)。它还能正确拼写 6 个未见过的手绘草图中的 4 个。相比之下,所有基线模型都无法一致地识别图像并正确拼写单词。对于较简单的灵巧任务,我们发现从头训练的单任务扩散模型具有竞争力,这与已发表的最佳结果一致。然而,为拼写游戏、折纸和午餐盒任务训练的单任务扩散模型表现不佳,可能是由于这些任务的长时程特性。我们还发现,多任务扩散和\(\pi_{0}\) 重实现模型在使用相同数据微调后,均未能达到我们模型的性能。这与我们在图 16 中的发现一致。Gemini Robotics 模型与基线模型之间的关键区别在于更强大的基于 Gemini 的主干网络,这表明在具有挑战性的任务上成功专门化与通用模型的强度高度相关。此外,当我们直接使用专门化数据集从头训练 Gemini Robotics 专门化模型时,我们发现它无法解决任何这些任务(成功率均为 0%,图 23 中未包含该图),这表明除了高容量模型架构外,从第 3 节中多样化的机器人动作数据集中学到的表示或物理常识,也是模型在需要高灵巧性的挑战性长时程任务中实现专门化的另一个关键组成部分。
We find that our specialist models can solve all these tasks with an average success rate of 79%. Most notably, it achieves a 100% success rate of the full long-horizon lunch-box packing task which takes over 2 minutes to complete. In the spelling game, it correctly reads and spells words from printed images (seen in the specialization dataset). It is also able to correctly spell 4 out of 6 unseen hand-drawn sketches. In contrast, none of the baselines can consistently recognize the images and spell the words correctly. For the simpler dexterous tasks, we find that the single task diffusion model that is trained from scratch is competitive, which is consistent with the best published results. However, the single task diffusion models trained for spelling game, origami, and lunch-box tasks perform poorly, possibly due to the long-horizon nature of these tasks. We also find that both Multi-task diffusion and \(\pi_{0}\) re-implement, after fine-tuning using the same data, fail to meet our model’s performance. This is consistent with our findings in Fig. 16. The key difference between the Gemini Robotics model and the baselines is the much more powerful Gemini-based backbone, which suggests that successful specialization on challenging tasks highly correlates with the strength of the generalist model. Furthermore, when we directly train the Gemini Robotics specialist model from scratch using the specialization datasets, we find that it is unable to solve any of these tasks (0% success rates across the board, and plot not included in Fig. 23), suggesting that in addition to the high-capacity model architecture, the representation, or the physical common sense, learned from diverse robot action datasets in Section 3 is another key component for the model to specialize in challenging long-horizon tasks that require a high level of dexterity.
我们现在探讨如何充分利用 Gemini Robotics-ER 带来的新型具身推理能力,例如空间与物理理解以及世界知识,来指导低层机器人动作,适用于需要推理且比第 3.4 节要求更广泛泛化的场景。尽管先前的工作在视觉鲁棒性方面取得了一致的提升,但迄今为止,VLA 在保持抽象推理能力并将其应用于行为泛化方面仍面临重大挑战。为此,我们研究了一种微调过程,该过程利用第 3.1 节中机器人动作数据集的重新标注版本,使动作预测更接近新引入的具身推理能力:轨迹理解与生成(第 2.2 节)。第 3.1 节中的局部动作解码器被扩展,以将这些推理中间结果转换为连续的低层动作。
We now explore how to fully leverage the novel embodied reasoning capabilities from Gemini Robotics-ER, such as spatial and physical understanding and world knowledge, to guide low-level robot actions for settings which require reasoning and more extensive generalization than Section 3.4. Although prior works have found consistent gains in visual robustness, so far VLAs still face substantial challenges in retaining abstract reasoning capabilities, and applying them to behavior generalization. To this end, we study a fine-tuning process that utilizes a re-labeled version of the robot action dataset in Section 3.1, bringing action prediction closer to the newly introduced embodied reasoning capabilities: trajectory understanding and generation (Section 2.2). The local action decoder from Section 3.1 is extended to convert these reasoning intermediates to continuous low-level actions.
我们将这种推理增强变体与原始 Gemini Robotics 模型(第 3 节)在训练分布之外的真实世界机器人任务(第 3.1 节)上进行比较。值得注意的是,这些具有挑战性的场景结合了第 3.4 节中研究的分布偏移,要求模型能够同时泛化到指令、视觉和动作变化。我们描述了高层评估类别,并在第 D.2 节中列出了完整的指令和任务描述。
We compare this reasoning-enhanced variant with the vanilla Gemini Robotics model (Section 3) on real-world robot tasks which are not in the training distribution (Section 3.1). Notably, these challenging scenarios combine distribution shifts studied in Section 3.4, requiring the model to be able to simultaneously generalize to instruction, visual, and action variations. We describe the high-level evaluation categories, and list the full instructions and task descriptions in Section D.2.
一步推理:对于此类任务,指令通过对象的属性或可供性等间接方式指定感兴趣的对象和/或操作动作。例如,在“将右下角的鼠标放入匹配的堆中”任务中,模型必须将右下角的白色玩具鼠标放入白色玩具鼠标堆中,而不是棕色和灰色玩具鼠标的干扰堆;所有这些鼠标以及基于颜色对物体进行分类的任务,在训练动作标签分布中都是未见过的。
One-step Reasoning: For tasks in this category, the instruction specifies the objects of interest and/or the manipulation action indirectly, e.g., via their properties or affordances. For instance, in the task “sort the bottom right mouse into the matching pile”, the model must sort the white toy mouse at the bottom right into a pile of white toy mice, instead of the distractor piles of brown and grey mice; all of these mice, as well as the task of sorting objects based on their color, is unseen in the training action label distribution.
语义泛化:这些任务需要超出第 3.4 节泛化任务复杂度的语义和视觉理解。对于“将日本鱼美食放入午餐盒”任务,模型必须在各种干扰物体中确定寿司是目标物体,并将寿司放入午餐盒。
Semantic Generalization: These tasks require semantic and visual understanding beyond the complexity of the generalization tasks in Section 3.4. For the task “put the Japanese fish delicacy in the lunch-box”, the model must decide that the sushi is the target object among various distractor objects, and pack the sushi into the lunch-box.
空间理解:这些任务需要理解相对和绝对空间关系的概念。对于“将最小的可乐罐放入午餐盒”任务,模型必须放入迷你罐而不是干扰的全尺寸罐,并将其放入午餐盒。描述所评估空间概念(最小)的语言在训练动作数据标签分布中是未见过的。
Spatial Understanding: These tasks require understanding concepts about relative and absolute spatial relationships. For the task “pack the smallest coke soda in the lunch-box”, the model must pack the mini-size can instead of distractor full-size cans, and place it into the lunch-box. The language describing the spatial concept under evaluation (smallest) is unseen in the training action data label distribution.
图 24 展示了原始 Gemini Robotics 模型及其推理增强版本在真实世界评估中的成功率。尽管原始模型的表现依然合理,但在需要单步推理或规划、语义知识以及空间理解的世界模型的分布外场景中,推理增强版本将成功率提升得更高。此外,除了模型在新环境中部署技能的能力提升之外,我们还观察到可解释性的增强,因为模型能够输出与 Gemini Robotics-ER 的人类可解释的具身推理轨迹高度相似的中间步骤,这一优势在先前启发性的工作中也得到了强调。作为示例,我们在图 25 中展示了关键点轨迹的可视化,这些轨迹被用作模型内部思维链的一部分。
Success rates of both the vanilla Gemini Robotics model and its reasoning-enhanced version in real-world evaluations are shown in Fig. 24. While the vanilla model still performs reasonably, the reasoning-enhanced version pushes the success rate much higher in out-of-distribution scenarios that require single-step reasoning or planning, semantic knowledge, and spatial understanding of the world. Additionally, beyond improvements in the model's ability to deploy its skills in novel settings, we also see increased interpretability as the model can output intermediate steps that closely resemble the human-interpretable embodied reasoning traces of Gemini Robotics-ER, a benefit also highlighted in inspiring prior works. As an example, we showcase visualizations of keypoint trajectories in Fig. 25, utilized as part of the model's internal chain of thought.
机器人基础模型有望通过利用预先获取的关于机器人动作和物理交互的常识来实现快速任务学习。第 4.1 节探讨了在长时程、高灵巧性任务上的专精化,而本节则研究光谱的另一端:我们的通用模型能够多快地适应新的、较短时程的任务。具体来说,我们从上述长时程任务中选取了八个子任务(详见第 D.3.1 节),并变化用于微调第 3 节检查点的数据量。图 26 展示了每个任务的平均成功率随演示次数的变化。在 8 个任务中,有 7 个任务在最多 100 次演示(根据任务复杂度,相当于 15 分钟到 1 小时的演示)下,微调能够有效实现超过 70%的成功率。值得一提的是,对于两个任务,Gemini Robotics 实现了 100%的成功率。基线方法在较简单的任务上具有竞争力:它们更高效地学会了“倒生菜”,而对于“沙拉酱”和“抽卡”,π0 重实现版本的成功率略高。然而,在更困难的任务(如“折纸狐狸第一步”或午餐盒任务)上,它们无法在有限的演示次数下表现良好。这再次证明,强大的 VLM 骨干网络能够更有效地将丰富多样的机器人动作数据转化为对物理交互的细致理解,是实现新任务快速学习的关键。
Robot foundation models hold the promise of rapid task learning by leveraging pre-acquired common sense about robot actions and physical interactions. While Section 4.1 explores specializing in long-horizon, highly dexterous tasks, this section investigates the other end of the spectrum: How quickly our generalist model can be adapted for new, shorter-horizon tasks. Concretely, we select eight sub-tasks (details in Section D.3.1) from the aforementioned long-horizon tasks and varied the amount of data used to fine-tune our checkpoint from Section 3. Fig. 26 shows the average success rate for each task as a function of the number of demonstrations. For 7 out of 8 tasks, fine-tuning was effective at achieving success rate above \(70\%\) with at most 100 demonstrations (equivalent to 15 minutes to 1 hour of demonstrations depending on the complexity of the task). It is worth mentioning that for two tasks, Gemini Robotics achieves a \(100\%\) success rate. Baselines are competitive on the easier tasks: they learn "Pour lettuce" more efficiently, and for "Salad dressing" and "Draw card", \(\pi_{0}\) re-implement achieves slightly higher success rate. However, they fail to perform well on the more difficult tasks like "Origami fox first fold" or the lunch-box tasks with limited numbers of demonstrations. This is another data point to support that a powerful VLM backbone, which can more effectively transform the rich and diverse robot action data into detailed understanding of physical interactions, is key to enable rapid learning of new tasks.
在初步实验中,我们还探索了如何将使用 ALOHA 2 上收集的动作数据训练的 Gemini Robotics 模型,通过目标平台上的少量数据高效地适应控制新的具身。我们考虑了带有平行夹爪的双臂 Franka 机器人,以及来自 Apptronik 的 Apollo——一个具有五指灵巧手的全尺寸人形机器人。图 27 展示了这两种不同机器人上的示例任务。经过微调后,我们发现 Gemini Robotics 在分布内任务上的成功率与最先进的单任务扩散策略相当或略优。例如,适应双臂 Franka 机器人的 Gemini Robotics 模型能够以平均成功率\(63\%\)解决所有考虑的任务(任务详情和成功率图见 D.4 节)。我们进一步研究了该适应模型对视觉干扰、初始条件扰动和物体形状变化的鲁棒性(D.4.2 节)。如图 28 所示,在这些视觉和动作泛化测试中,Gemini Robotics 显著优于单任务扩散基线。值得注意的是,这表明 Gemini Robotics 模型能够将其鲁棒性和泛化能力跨不同具身进行迁移,即使经过针对新具身的微调后也是如此。
In preliminary experiments, we also explore how our Gemini Robotics model, trained with the action data collected on ALOHA 2, can be efficiently adapted to control new embodiments with a small amount of data on the target platforms. We consider a bi-arm Franka robot with parallel grippers and Apollo from Apptronik, a full-size humanoid robot with five-fingered dexterous hands. Fig. 27 shows example tasks on these two different robots. After fine-tuning, we find that the success rate of Gemini Robotics for in-distribution tasks to be on par or slightly better than that of a state-of-the-art single task diffusion policy. For instance, the adapted Gemini Robotics model for the bi-arm Franka robot can solve all considered tasks with an average success rate of \(63\%\) (tasks details and plots of success rate available in Section D.4). We further investigate the robustness of this adapted model to visual disturbances, initial condition perturbations, and object shape variations (Section D.4.2). As illustrated in Fig. 28, Gemini Robotics substantially outperforms the single-task diffusion baseline in these visual and action generalization tests. Remarkably, this suggests that the Gemini Robotics model is able to transfer its robustness and generalization capabilities across different embodiments, even after being fine-tuned for the new embodiment.
我们开发本报告中的模型时,遵循了谷歌 AI 原则和先前 AI 技术的发布规范。确保 AI 的负责任构建和使用是一个迭代过程——这同样适用于机器人基础模型,也适用于文本或图像模型。我们模型的混合数字-物理和具身特性,以及它们最终使机器人能够在物理世界中行动的事实,需要特别考虑。在谷歌 DeepMind 的责任与安全委员会(RSC)和负责任开发与创新(ReDI)团队的指导下,我们识别了使用我们模型的风险,并制定了安全缓解框架,以覆盖我们模型的具身推理和动作输出模态。
We have developed the models introduced in this report in alignment with Google AI Principles and previous releases of AI technology. Ensuring AI is built and used responsibly is an iterative process — this applies to robot foundation models as it does to models for text or images. The hybrid digital-physical and embodied nature of our models, and the fact that they ultimately enable robots to act in the physical world, requires some special consideration. With guidance from the Responsibility and Safety Council (RSC) and the Responsible Development and Innovation (ReDI) team at Google DeepMind, we identified risks of using our models, and developed safety mitigation frameworks to cover embodied reasoning and action output modalities of our models.
传统机器人安全是一个广泛而多方面的学科,从数百页 ISO 和 RIA 标准中规定的危害缓解,到无碰撞运动规划、力调制和鲁棒控制。历史上,重点一直放在物理动作安全上,即确保机器人遵守硬性物理约束(如避障、工作空间边界)、具有稳定的移动性(如运动控制),并能将接触力调节在安全范围内。这属于经典约束控制的范畴,在控制栈的最低层通过运动规划、模型预测控制和柔顺/力控制等方法实现。根据硬件特性和环境约束,我们需要将 Gemini Robotics 等 VLA 模型与此类安全关键的低层控制器接口。我们先前的研究已经原型化了此类接口。此外,本报告描述的 AI 驱动机器人系统类别需要更广泛且不断发展的安全研究视角,因为新的安全概念变得相关。
Traditional robot safety is a vast multifaceted discipline ranging from hazard mitigation codified in hundreds of pages of ISO and RIA standards, to collision-free motion planning, force modulation and robust control. Historically, the focus has been on physical action safety, i.e., on ensuring that robots respect hard physical constraints (e.g., obstacle avoidance, workspace bounds), have stable mobility (e.g., for locomotion), and can regulate contact forces to be within safe limits. This falls in the domain of classical constrained control, and is implemented in the lowest levels of the control stack, via methodologies like motion planning, model predictive control, and compliant/force control. Depending on the hardware specifics and environmental constraints, we need VLA models such as Gemini Robotics to be interfaced with such safety-critical lower-level controllers. Our prior research has prototyped such interfaces. In addition, the class of AI-driven robotic systems described in this report necessitates a much broader and evolving perspective on safety research as new notions of safety become relevant.
Gemini 安全政策旨在确保内容安全,防止源自 Gemini 的模型生成有害的对话内容,如仇恨言论、色情内容、不当医疗建议以及泄露个人身份信息。通过基于 Gemini 检查点,我们的机器人模型继承了这些政策的安全训练,促进了安全的人机对话。由于我们的具身推理模型引入了新的输出模态,如指向,我们需要为这些新功能增加额外的内容安全层。因此,我们对 Gemini 2.0 和 Gemini Robotics-ER 进行了监督微调,目的是教会 Gemini 在何时应用超出图像可用信息的泛化是不合适的。这种训练使得对引发偏见的指向查询的拒绝率达到 96%,而基线率为 20%。
Gemini Safety policies outlined in are designed for content safety, preventing Gemini-derived models from generating harmful conversational content such as hate speech, sexual explicitness, improper medical advice, and revealing personally identifiable information. By building on Gemini checkpoints, our robotics models inherit safety training for these policies done in , promoting safe human-robot dialog. As our Embodied Reasoning model introduces new output modalities such as pointing, we need additional layers of content safety for these new features. We therefore perform supervised fine-tuning on both Gemini 2.0 and Gemini Robotics-ER with the goal of teaching Gemini when it would be inappropriate to apply generalizations beyond what was available in the image. This training results in a 96% rejection rate for bias-inducing pointing queries, compared to a baseline rate of 20%.
除了内容安全,通用机器人一个重要考虑是语义动作安全,即在开放域非结构化环境中遵守物理安全约束的需要。这些约束难以穷举——例如,软玩具不能放在热炉子上;过敏者不能吃花生;酒杯必须保持直立转移;刀不能指向人;等等。这些考虑不仅适用于通用机器人,也适用于其他情境智能体。与本技术报告同时,我们开发并发布了 ASIMOV 数据集,以评估和改进语义动作安全。该数据包括视觉和纯文本的安全问答实例,如图 29(a)和图 29(b)所示。Gemini Robotics-ER 模型在这些实例上进行了后训练。我们的安全评估总结在图 29(c)和 29(d)中。对齐度量是相对于真实人类安全评估的二元分类准确率。我们在图 29(c)和 29(d)中看到,Gemini 2.0 Flash 和 Gemini Robotics-ER 模型表现相似,分别展示了在视觉场景和来自真实世界伤害报告的场景中对物理安全的强大语义理解。我们使用宪法 AI 方法看到了性能改进。我们还看到,在对抗性提示(要求模型翻转其对可取和不可取的理解)下的性能下降可以通过后训练和宪法 AI 机制来缓解。关于 ASIMOV 基准、我们的数据驱动宪法生成过程以及全面实证分析的更多细节,请参见与本技术报告同时发布的文档。
Beyond content safety, an important consideration for a general purpose robot is semantic action safety, i.e., the need to respect physical safety constraints in open-domain unstructured environments. These are hard to exhaustively enumerate – that a soft toy must not be placed on a hot stove; an allergic person must not be served peanuts; a wine glass must be transferred in upright orientation; a knife should not be pointed at a human; and so on. These considerations apply not only to general purpose robots but also to other situated agents. Concurrent with this tech report, we develop and release the ASIMOV-datasets to evaluate and improve semantic action safety. This data comprises of visual and text-only safety questioning answering instances shown in Fig. 29(a) and Fig. 29(b). Gemini Robotics-ER models are post-trained on such instances. Our safety evaluations are summarized in Fig. 29(c) and 29(d). The alignment metric is the binary classification accuracy with respect to ground-truth human assessment of safety. We see in Fig. 29(c) and 29(d) that both Gemini 2.0 Flash and Gemini Robotics-ER models perform similarly, demonstrating strong semantic understanding of physical safety in visual scenes and scenarios drawn from real-world injury reports respectively. We see performance improvements with the use of constitutional AI methods. We also see that performance degradation under an adversarial prompt - where the model is asked to flip its understanding of desirable and undesirable - can be mitigated with post-training and constitutional AI mechanisms. For more details on the ASIMOV benchmark, our data-driven constitution generation process, and comprehensive empirical analysis, see released concurrently with this tech report.
这些调查提供了一些初步保证,即我们非机器人模型所坚持的严格安全标准也适用于我们新的具身和机器人聚焦模型。随着我们进一步发展机器人基础模型家族,我们将继续改进和创新安全和对齐方法。除了潜在的安全风险,我们还必须承认机器人部署的社会影响。我们相信,主动监测和管理这些影响,包括收益和挑战,对于风险缓解、负责任部署和透明报告至关重要。Gemini Robotics 模型的模型卡可在附录 A 中找到。
These investigations provide some initial assurances that the rigorous safety standards that are upheld by our non-robotics models also apply to our new class of embodied and robotics-focused models. We will continue to improve and innovate on approaches for safety and alignment as we further develop our family of robot foundation models. Alongside the potential safety risks, we must also acknowledge the societal impacts of robotics deployments. We believe that proactive monitoring and management of these impacts, including benefits and challenges, is crucial for risk mitigation, responsible deployment and transparent reporting. The model card for Gemini Robotics models can be found in Appendix A.
在这项工作中,我们研究了如何通过机器人技术将 Gemini 2.0 的世界知识和推理能力带入物理世界。稳健的类人具身推理对于机器人和其他物理实体智能体至关重要。为此,我们推出了 Gemini Robotics-ER,一个具身视觉语言模型(VLM),它在空间理解、轨迹预测、多视角对应和精确指向方面显著推进了现有技术水平。我们通过一个新的开源基准验证了 Gemini Robotics-ER 的强劲性能。结果表明,我们的训练流程在增强 Gemini 2.0 固有的多模态能力以用于具身推理方面非常有效。所得模型为现实世界的机器人应用奠定了坚实基础,能够高效地实现零样本和少样本适应,用于感知、规划和控制机器人的代码生成等任务。
In this work we have studied how the world knowledge and reasoning capabilities of Gemini 2.0 can be brought into the physical world through robotics. Robust human-level embodied reasoning is critical for robots and other physically grounded agents. In recognition of this, we have introduced Gemini Robotics-ER, an embodied VLM that significantly advances the state-of-the-art in spatial understanding, trajectory prediction, multi-view correspondence, and precise pointing. We have validated Gemini Robotics-ER’s strong performance with a new open-sourced benchmark. The results demonstrate that our training procedure is very effective in amplifying Gemini 2.0’s inherent multimodal capabilities for embodied reasoning. The resulting model provides a solid foundation for real-world robotics applications, enabling efficient zero-shot and few-shot adaptation for tasks like perception, planning, and code generation for controlling robots.
我们还介绍了 Gemini Robotics,一个通用的视觉-语言-动作模型(VLA),它建立在 Gemini Robotics-ER 的基础上,弥合了被动感知与主动具身交互之间的鸿沟。作为我们迄今为止最灵巧的通用模型,Gemini Robotics 在多种操作任务中表现出色,从复杂的布料操作到对铰接物体的精确处理。我们推测,我们方法的成功可归因于:(1)具备增强具身推理能力的强大视觉语言模型;(2)我们针对机器人领域的特定训练方案,该方案结合了海量机器人动作数据和多样化的非机器人数据;(3)其专为低延迟机器人控制设计的独特架构。至关重要的是,Gemini Robotics 能有效遵循开放词汇指令,并展现出强大的零样本泛化能力,证明了其能够利用 Gemini Robotics-ER 的具身推理能力。最后,我们展示了可选的微调以实现专业化和适应,使 Gemini Robotics 能够适应新任务和新具身形态,实现极致灵巧性,并在具有挑战性的场景中泛化,从而凸显了我们方法在将基础能力快速转化为实际应用方面的灵活性和实用性。
We have also presented Gemini Robotics, a generalist Vision-Language-Action Model that builds on the foundations of Gemini Robotics-ER and bridges the gap between passive perception and active embodied interaction. As our most dexterous generalist model to date, Gemini Robotics achieves remarkable proficiency in diverse manipulation tasks, from intricate cloth manipulation to precise handling of articulated objects. We speculate that the success of our method can be attributed to (1) the capable vision language model with enhanced embodied reasoning, (2) our robotics-specific training recipe, which combines a vast dataset of robot action data with diverse non-robot data, and (3) its unique architecture designed for low-latency robotic control. Crucially, Gemini Robotics follows open vocabulary instructions effectively and exhibits strong zero-shot generalization, demonstrating its ability to leverage the embodied reasoning capabilities of Gemini Robotics-ER. Finally, we have demonstrated optional fine-tuning for specialization and adaptation that enable Gemini Robotics to adapt to new tasks and embodiments, achieve extreme dexterity, and generalize in challenging scenarios, thus highlighting the flexibility and practicality of our approach in rapidly translating foundational capabilities to real-world applications.
局限性与未来工作。Gemini 2.0 和 Gemini Robotics-ER 在具身推理方面取得了显著进展,但其能力仍有提升空间。例如,Gemini 2.0 在处理长视频中的空间关系接地时可能遇到困难,其数值预测(如点和框)可能不够精确,难以满足更细粒度的机器人控制任务。此外,虽然我们使用 Gemini Robotics 的初步结果展示了有前景的泛化能力,但未来的工作将集中在几个关键领域。首先,我们旨在增强 Gemini Robotics 处理需要多步推理和精确灵巧动作的复杂场景的能力,尤其是在新情境中。这涉及开发技术,将抽象推理与精确执行无缝集成,从而实现更稳健和更可泛化的性能。其次,我们计划更多地利用仿真来生成视觉多样且接触丰富的数据,并开发技术利用这些数据构建更强大的 VLA 模型,使其能够迁移到现实世界。最后,我们将扩展多具身实验,旨在减少适应新型机器人所需的数据,最终实现零样本跨具身迁移,使模型能够立即将其技能泛化到新的机器人平台。
Limitations and future work. Gemini 2.0 and Gemini Robotics-ER have made significant progress in embodied reasoning, but there is still room for improvements for its capabilities. For example, Gemini 2.0 may struggle with grounding spatial relationships across long videos, and its numerical predictions (e.g., points and boxes) may not be precise enough for more fine-grained robot control tasks. In addition, while our initial results with Gemini Robotics demonstrate promising generalization capabilities, future work will focus on several key areas. First, we aim to enhance Gemini Robotics’s ability to handle complex scenarios requiring both multi-step reasoning and precise dexterous movements, particularly in novel situations. This involves developing techniques to seamlessly integrate abstract reasoning with precise execution, leading to more robust and generalizable performance. Second, we plan to lean more on simulation to generate visually diverse and contact rich data as well as developing techniques for using this data to build more capable VLA models that can transfer to the real world. Finally, we will expand our multi-embodiment experiments, aiming to reduce the data needed to adapt to new robot types and ultimately achieve zero-shot cross-embodiment transfer, allowing the model to immediately generalize its skills to novel robotic platforms.
总之,我们的工作朝着实现物理世界中通用自主 AI 的愿景迈出了重要一步。这将带来机器人系统理解、学习和被指令方式的范式转变。传统机器人系统是为特定任务而构建的,而 Gemini Robotics 为机器人提供了对世界运作方式的通用理解,使其能够适应广泛的任务。Gemini 的多模态、通用特性进一步有可能降低使用和受益于机器人技术的技术门槛。未来,这可能从根本上改变机器人系统的应用场景和使用者,最终使智能机器人能够部署到我们的日常生活中。因此,随着技术的成熟,像 Gemini Robotics 这样强大的机器人模型将具有巨大的潜力,为社会带来积极影响。但同样重要的是要考虑其安全性和更广泛的社会影响。Gemini Robotics 在设计时考虑了安全性,我们已讨论了几种缓解策略。未来,我们将继续努力确保这些技术的潜力得到安全、负责任地利用。
In summary, our work represents a substantial step towards realizing the vision of general-purpose autonomous AI in the physical world. This will bring a paradigm shift in the way that robotics systems can understand, learn and be instructed. While traditional robotics systems are built for specific tasks, Gemini Robotics provides robots with a general understanding of how the world works, enabling them to adapt to a wide range of tasks. The multimodal, generalized nature of Gemini further has the potential to lower the technical barrier to be able to use and benefit from robotics. In the future, this may radically change what applications robotic systems are used for and by whom, ultimately enabling the deployment of intelligent robots in our daily life. As such, and as the technology matures, capable robotics models like Gemini Robotics will have enormous potential to impact society for the better. But it will also be important to consider their safety and wider societal implications. Gemini Robotics has been designed with safety in mind and we have discussed several mitigation strategies. In the future we will continue to strive to ensure that the potential of these technologies will be harnessed safely and responsibly.
作者:Saminda Abeyruwan、Joshua Ainslie、Jean-Baptiste Alayrac、Montserrat Gonzalez Arenas、Travis Armstrong、Ashwin Balakrishna、Robert Baruch、Maria Bauza、Michiel Blokzijl、Steven Bohez、Konstantinos Bousmalis、Anthony Brohan、Thomas Buschmann、Arunkumar Byravan、Serkan Cabi、Ken Caluwaerts、Federico Casarini、Oscar Chang、Jose Enrique Chen、Xi Chen、Hao-Tien Lewis Chiang、Krzysztof Choromanski、David D’Ambrosio、Sudeep Dasari、Todor Davchev、Coline Devin、Norman Di Palo、Tianli Ding、Adil Dostmohamed、Danny Driess、Yilun Du、Debidatta Dwibedi、Michael Elabd、Claudio Fantacci、Cody Fong、Erik Frey、Chuyuan Fu、Marissa Giustina、Keerthana Gopalakrishnan、Laura Graesser、Leonard Hasenclever、Nicolas Heess、Brandon Hernaez、Alexander Herzog、R. Alex Hofer、Jan Humplik、Atil Iscen、Mithun George Jacob、Deepali Jain、Ryan Julian、Dmitry Kalashnikov、M. Emre Karagozler、Stefani Karp、Chase Kew、Jerad Kirkland、Sean Kirmani、Yuheng Kuang、Thomas Lampe、Antoine Laurens、Isabel Leal、Alex X. Lee、Tsang-Wei Edward Lee、Jacky Liang、Yixin Lin、Sharath Maddineni、Anirudha Majumdar、Assaf Hurwitz Michaely、Robert Moreno、Michael Neunert、Francesco Nori、Carolina Parada、Emilio Parisotto、Peter Pastor、Acorn Pooley、Kanishka Rao、Krista Reymann、Dorsa Sadigh、Stefano Saliceti、Pannag Sanketi、Pierre Sermanet、Dhruv Shah、Mohit Sharma、Kathryn Shea、Charles Shu、Vikas Sindhwani、Sumeet Singh、Radu Soricut、Jost Tobias Springenberg、Rachel Sterneck、Razvan Surdulescu、Jie Tan、Jonathan Tompson、Vincent Vanhoucke、Jake Varley、Grace Vesom、Giulia Vezzani、Oriol Vinyals、Ayzaan Wahid、Stefan Welker、Paul Wohlhart、Fei Xia、Ted Xiao、Annie Xie、Jinyu Xie、Peng Xu、Sichun Xu、Ying Xu、Zhuo Xu、Yuxiang Yang、Rui Yao、Sergey Yaroshenko、Wenhao Yu、Wentao Yuan、Jingwei Zhang、Tingnan Zhang、Allan Zhou、Yuxiang Zhou。
Authors: Saminda Abeyruwan, Joshua Ainslie, Jean-Baptiste Alayrac, Montserrat Gonzalez Arenas, Travis Armstrong, Ashwin Balakrishna, Robert Baruch, Maria Bauza, Michiel Blokzijl, Steven Bohez, Konstantinos Bousmalis, Anthony Brohan, Thomas Buschmann, Arunkumar Byravan, Serkan Cabi, Ken Caluwaerts, Federico Casarini, Oscar Chang, Jose Enrique Chen, Xi Chen, Hao-Tien Lewis Chiang, Krzysztof Choromanski, David D’Ambrosio, Sudeep Dasari, Todor Davchev, Coline Devin, Norman Di Palo, Tianli Ding, Adil Dostmohamed, Danny Driess, Yilun Du, Debidatta Dwibedi, Michael Elabd, Claudio Fantacci, Cody Fong, Erik Frey, Chuyuan Fu, Marissa Giustina, Keerthana Gopalakrishnan, Laura Graesser, Leonard Hasenclever, Nicolas Heess, Brandon Hernaez, Alexander Herzog, R. Alex Hofer, Jan Humplik, Atil Iscen, Mithun George Jacob, Deepali Jain, Ryan Julian, Dmitry Kalashnikov, M. Emre Karagozler, Stefani Karp, Chase Kew, Jerad Kirkland, Sean Kirmani, Yuheng Kuang, Thomas Lampe, Antoine Laurens, Isabel Leal, Alex X. Lee, Tsang-Wei Edward Lee, Jacky Liang, Yixin Lin, Sharath Maddineni, Anirudha Majumdar, Assaf Hurwitz Michaely, Robert Moreno, Michael Neunert, Francesco Nori, Carolina Parada, Emilio Parisotto, Peter Pastor, Acorn Pooley, Kanishka Rao, Krista Reymann, Dorsa Sadigh, Stefano Saliceti, Pannag Sanketi, Pierre Sermanet, Dhruv Shah, Mohit Sharma, Kathryn Shea, Charles Shu, Vikas Sindhwani, Sumeet Singh, Radu Soricut, Jost Tobias Springenberg, Rachel Sterneck, Razvan Surdulescu, Jie Tan, Jonathan Tompson, Vincent Vanhoucke, Jake Varley, Grace Vesom, Giulia Vezzani, Oriol Vinyals, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Fei Xia, Ted Xiao, Annie Xie, Jinyu Xie, Peng Xu, Sichun Xu, Ying Xu, Zhuo Xu, Yuxiang Yang, Rui Yao, Sergey Yaroshenko, Wenhao Yu, Wentao Yuan, Jingwei Zhang, Tingnan Zhang, Allan Zhou, Yuxiang Zhou.
致谢:我们的工作得益于谷歌众多团队的奉献和努力。我们要感谢以下人员的支持:Adrian Collister、Alan Thompson、Alessio Quaglino、Anca Dragan、Ashley Gibb、Ben Bariach、Caden Lu、Catarina Barros、Christine Chan、Clara Barbu、Dave Orr、Demetra Brady、Dhruva Tirumala、Dushyant Rao、Francesco Romano、Frankie Garcia、Grace Popple、Haroon Qureshi、Howard Zhou、Huizhong Chen、Jennie Lees、Joss Moore、Karen Truong、Kendra Byrne、Keran Rong、Kevis-Kokitsi Maninis、Kieran Connell、Markus Wulfmeier、Martina Zambelli、Matt Young、Mili Sanwalka、Mohit Shridhar、Nathan Batchelor、Sally Jesmonth、Sam Haves、Sandy H Huang、Simon Green、Siobhan Mcloughlin、Tom Erez、Yanan Bao、Yuval Tassa 和 Zhicheng Wang。
Acknowledgements: Our work is made possible by the dedication and efforts of numerous teams at Google. We would like to acknowledge the support from Adrian Collister, Alan Thompson, Alessio Quaglino, Anca Dragan, Ashley Gibb, Ben Bariach, Caden Lu, Catarina Barros, Christine Chan, Clara Barbu, Dave Orr, Demetra Brady, Dhruva Tirumala, Dushyant Rao, Francesco Romano, Frankie Garcia, Grace Popple, Haroon Qureshi, Howard Zhou, Huizhong Chen, Jennie Lees, Joss Moore, Karen Truong, Kendra Byrne, Keran Rong, Kevis-Kokitsi Maninis, Kieran Connell, Markus Wulfmeier, Martina Zambelli, Matt Young, Mili Sanwalka, Mohit Shridhar, Nathan Batchelor, Sally Jesmonth, Sam Haves, Sandy H Huang, Simon Green, Siobhan Mcloughlin, Tom Erez, Yanan Bao, Yuval Tassa and Zhicheng Wang.
我们还要感谢谷歌和谷歌 DeepMind 的众多团队,包括谷歌创意实验室、法律、市场营销、传播、责任与安全委员会、负责任发展与创新、政策、战略与运营以及我们的业务和企业发展团队。我们要感谢机器人团队中所有未在上文明确提及的成员,感谢他们持续的支持和指导。我们还要感谢 Apptronik 团队的支持。
We would also like to recognize the many teams across Google and Google DeepMind that have contributed to this effort including Google Creative Lab, Legal, Marketing, Communications, Responsibility and Safety Council, Responsible Development and Innovation, Policy, Strategy and Operations as well as our Business and Corporate Development teams. We would like to thank everyone on the Robotics team not explicitly mentioned above for their continued support and guidance. We would also like to thank the Apptronik team for their support.