照我做:从日常人类视频中获取灵巧操作数据

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos

彼得·阿贝尔 Pieter Abbeel · UC Berkeley · 2026-06-17 · arXiv:2606.19333 ↗ · 被引 2

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

如何可扩展地生成机器人操作数据,尤其是在灵巧多指手等类人平台上?从人类视频中学习近来成为这一问题的可能答案。然而,估计手-物体交互以及跨越人类到机器人的具身差距的困难,阻碍了将丰富的单目 RGB 人类视频作为机器人操作数据主要来源的采用。在这项工作中,我们提出了 DO AS I DO,一种从单目 RGB 人类视频重建并重定向到多指灵巧机器人手的算法。DO AS I DO 从各种自我中心和非自我中心的野外视频源中重建手-物体交互。然后,该算法将这些手-物体交互估计重定向为可在现实世界中执行的动作序列,从不同的人类视频中产生机器人完整的操作数据。总体而言,DO AS I DO 在估计手-物体交互和从 RGB 视频中提取灵巧操作轨迹方面优于先前的最先进技术,我们在具有真实值的数据集以及在线收集的视频片段数据集上的实验证明了这一点。我们的实验使我们能够为收集人类操作数据的实践者提出一个效能指南。

How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present DO AS I DO, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. DO AS I DO reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, DO AS I DO outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show in experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 12)

阅读逐段中英对照全文 →