$π_{0.5}$:一个具有开放世界泛化能力的视觉-语言-动作模型

$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

凯文·布莱克 Kevin Black · Physical Intelligence · 2025-04-22 · arXiv:2504.16054 ↗ · 被引 1288

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

为了让机器人真正有用,它们必须在实验室外的现实世界中执行实际相关的任务。尽管视觉-语言-动作(VLA)模型在端到端机器人控制方面取得了令人印象深刻的结果,但这些模型在野外能泛化到何种程度仍是一个开放问题。我们描述了$π_{0.5}$,这是一个基于$π_{0}$的新模型,通过异构任务的协同训练实现了广泛的泛化。$π_{0.5}$利用来自多个机器人的数据、高层语义预测、网络数据和其他来源,实现了广泛泛化的现实世界机器人操作。我们的系统结合了协同训练和混合多模态示例,这些示例结合了图像观察、语言指令、物体检测、语义子任务预测和低级动作。我们的实验表明,这种知识迁移对于有效泛化至关重要,并且我们首次证明,一个端到端学习驱动的机器人系统可以在全新的家庭环境中执行长时程和灵巧的操作技能,例如清洁厨房或卧室。

In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated impressive results for end-to-end robot control, it remains an open question how far such models can generalize in the wild. We describe $π_{0.5}$, a new model based on $π_{0}$ that uses co-training on heterogeneous tasks to enable broad generalization. $π_{0.5}$\ uses data from multiple robots, high-level semantic prediction, web data, and other sources to enable broadly generalizable real-world robotic manipulation. Our system uses a combination of co-training and hybrid multi-modal examples that combine image observations, language commands, object detections, semantic subtask prediction, and low-level actions. Our experiments show that this kind of knowledge transfer is essential for effective generalization, and we demonstrate for the first time that an end-to-end learning-enabled robotic system can perform long-horizon and dexterous manipulation skills, such as cleaning a kitchen or bedroom, in entirely new homes.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 18)

阅读逐段中英对照全文 →