GR00T N1:面向通用人形机器人的开放基础模型

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

范麟熙 Jim Fan · NVIDIA · 2025-03-18 · arXiv:2503.14734 ↗ · 被引 1007

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

通用机器人需要多功能的身体和智能的大脑。人形机器人作为在人类世界中构建通用自主性的硬件平台,近期展现出巨大潜力。机器人基础模型需基于海量多样化数据训练,使其能推理新场景、稳健应对现实变化并快速学习新任务。为此,我们提出 GR00T N1——一个开放的人形机器人基础模型。GR00T N1 是一种视觉-语言-动作(VLA)模型,采用双系统架构:视觉语言模块(系统 2)通过视觉和语言指令解读环境,扩散变换器模块(系统 1)实时生成流畅的电机动作。两个模块紧密耦合,端到端联合训练。我们使用真实机器人轨迹、人类视频和合成数据集的异构混合训练 GR00T N1。实验表明,我们的通用机器人模型 GR00T N1 在多个机器人形态的标准仿真基准上优于最先进的模仿学习基线。此外,我们将模型部署在 Fourier GR-1 人形机器人上执行语言条件双臂操作任务,以高数据效率实现了强劲性能。

General-purpose robots need a versatile body and an intelligent mind. Recent advancements in humanoid robots have shown great promise as a hardware platform for building generalist autonomy in the human world. A robot foundation model, trained on massive and diverse data sources, is essential for enabling the robots to reason about novel situations, robustly handle real-world variability, and rapidly learn new tasks. To this end, we introduce GR00T N1, an open foundation model for humanoid robots. GR00T N1 is a Vision-Language-Action (VLA) model with a dual-system architecture. The vision-language module (System 2) interprets the environment through vision and language instructions. The subsequent diffusion transformer module (System 1) generates fluid motor actions in real time. Both modules are tightly coupled and jointly trained end-to-end. We train GR00T N1 with a heterogeneous mixture of real-robot trajectories, human videos, and synthetically generated datasets. We show that our generalist robot model GR00T N1 outperforms the state-of-the-art imitation learning baselines on standard simulation benchmarks across multiple robot embodiments. Furthermore, we deploy our model on the Fourier GR-1 humanoid robot for language-conditioned bimanual manipulation tasks, achieving strong performance with high data efficiency.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 28)

阅读逐段中英对照全文 →