$π_0$: 面向通用机器人控制的视觉-语言-动作流模型

$π_0$: A Vision-Language-Action Flow Model for General Robot Control

凯文·布莱克 Kevin Black · Physical Intelligence · 2024-10-31 · arXiv:2410.24164 ↗ · 被引 2092

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

机器人学习有望释放灵活、通用、灵巧机器人系统的全部潜力,并解决人工智能领域最深层次的问题。然而,要使机器人学习达到实际系统所需的通用性水平,在数据、泛化能力和鲁棒性方面面临重大障碍。本文探讨了通用机器人策略(即机器人基础模型)如何应对这些挑战,以及如何为复杂且高度灵巧的任务设计有效的通用机器人策略。我们提出了一种基于预训练视觉-语言模型(VLM)的新型流匹配架构,以继承互联网规模的语义知识。随后,我们讨论了如何利用来自多个灵巧机器人平台(包括单臂机器人、双臂机器人和移动操作器)的大规模多样化数据集训练该模型。我们从零样本预训练任务执行能力、遵循人类及高级 VLM 策略的语言指令能力,以及通过微调获取新技能的能力等方面评估了模型。我们的结果涵盖了多种任务,如叠衣服、清洁桌子和组装盒子。

Robot learning holds tremendous promise to unlock the full potential of flexible, general, and dexterous robot systems, as well as to address some of the deepest questions in artificial intelligence. However, bringing robot learning to the level of generality required for effective real-world systems faces major obstacles in terms of data, generalization, and robustness. In this paper, we discuss how generalist robot policies (i.e., robot foundation models) can address these challenges, and how we can design effective generalist robot policies for complex and highly dexterous tasks. We propose a novel flow matching architecture built on top of a pre-trained vision-language model (VLM) to inherit Internet-scale semantic knowledge. We then discuss how this model can be trained on a large and diverse dataset from multiple dexterous robot platforms, including single-arm robots, dual-arm robots, and mobile manipulators. We evaluate our model in terms of its ability to perform tasks in zero shot after pre-training, follow language instructions from people and from a high-level VLM policy, and its ability to acquire new skills via fine-tuning. Our results cover a wide variety of tasks, such as laundry folding, table cleaning, and assembling boxes.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →