GR-3 技术报告

GR-3 Technical Report

李星航 Xinghang Li · ByteDance Seed · 2025-07-21 · arXiv:2507.15493 ↗ · 被引 102

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们报告了在构建通用机器人策略方面的最新进展,即 GR-3 的开发。GR-3 是一个大规模视觉-语言-动作(VLA)模型。它在泛化到新物体、新环境以及涉及抽象概念的指令方面展现出卓越的能力。此外,它可以通过最少的人类轨迹数据进行高效微调,从而快速且经济地适应新环境。GR-3 还擅长处理长程和灵巧任务,包括需要双臂操作和移动的任务,展现出稳健可靠的性能。这些能力是通过多方面的训练方案实现的,包括与网络规模的视觉-语言数据共同训练、通过 VR 设备收集的人类轨迹数据进行高效微调,以及利用机器人轨迹数据进行有效的模仿学习。此外,我们介绍了 ByteMini,一种多功能双臂移动机器人,具有出色的灵活性和可靠性,在与 GR-3 集成后能够完成广泛的任务。通过大量的真实世界实验,我们展示了 GR-3 在各种具有挑战性的任务上超越了最先进的基线方法 $π_0$。我们希望 GR-3 能够成为构建能够协助人类日常生活的通用机器人的一步。

We report our recent progress towards building generalist robot policies, the development of GR-3. GR-3 is a large-scale vision-language-action (VLA) model. It showcases exceptional capabilities in generalizing to novel objects, environments, and instructions involving abstract concepts. Furthermore, it can be efficiently fine-tuned with minimal human trajectory data, enabling rapid and cost-effective adaptation to new settings. GR-3 also excels in handling long-horizon and dexterous tasks, including those requiring bi-manual manipulation and mobile movement, showcasing robust and reliable performance. These capabilities are achieved through a multi-faceted training recipe that includes co-training with web-scale vision-language data, efficient fine-tuning from human trajectory data collected via VR devices, and effective imitation learning with robot trajectory data. In addition, we introduce ByteMini, a versatile bi-manual mobile robot designed with exceptional flexibility and reliability, capable of accomplishing a wide range of tasks when integrated with GR-3. Through extensive real-world experiments, we show GR-3 surpasses the state-of-the-art baseline method, $π_0$, on a wide variety of challenging tasks. We hope GR-3 can serve as a step towards building generalist robots capable of assisting humans in daily life.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →