如何可扩展地生成机器人操作数据,尤其是在灵巧多指手等类人平台上?从人类视频中学习近来成为这一问题的可能答案。然而,估计手-物体交互以及跨越人类到机器人的具身差距的困难,阻碍了将丰富的单目 RGB 人类视频作为机器人操作数据主要来源的采用。在这项工作中,我们提出了 DO AS I DO,一种从单目 RGB 人类视频重建并重定向到多指灵巧机器人手的算法。DO AS I DO 从各种自我中心和非自我中心的野外视频源中重建手-物体交互。然后,该算法将这些手-物体交互估计重定向为可在现实世界中执行的动作序列,从不同的人类视频中产生机器人完整的操作数据。总体而言,DO AS I DO 在估计手-物体交互和从 RGB 视频中提取灵巧操作轨迹方面优于先前的最先进技术,我们在具有真实值的数据集以及在线收集的视频片段数据集上的实验证明了这一点。我们的实验使我们能够为收集人类操作数据的实践者提出一个效能指南。
How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. However, difficulties in estimating hand-object interaction and crossing the human-to-robot embodiment gap have hindered the adoption of abundant monocular RGB-only human videos as the primary source of robot manipulation data. In this work, we present DO AS I DO, an algorithm to reconstruct and retarget monocular RGB human videos to multi-fingered dexterous robotic hands. DO AS I DO reconstructs hand-object interactions from various egocentric and exocentric in-the-wild video sources. The algorithm then retargets these hand-object interaction estimates into a sequence of actions executable in the real world, yielding robot-complete manipulation data from disparate human videos. Overall, DO AS I DO outperforms previous state of the art in estimating hand-object interactions and extracting dexterous manipulation trajectories from RGB videos, as we show in experiments on datasets with ground truths and on a dataset of video clips collected online. Our experiments enable us to propose an efficacy playbook for practitioners collecting human data for manipulation.
核心贡献 · Key contributions
提出 Do as I Do,一种从单目 RGB 视频重建并重定向行为到多指灵巧手的两步算法。 Introduces Do as I Do, a two-step algorithm to reconstruct and retarget behaviors from monocular RGB videos to multi-fingered dexterous hands.
手-物体重建在相关指标上优于现有技术,并处理多样化的视频源。 Hand-object reconstruction outperforms SOTA on relevant metrics and handles diverse video sources.
重定向通过新颖组件改进了现有可扩展动力学感知技术,适用于噪声参考。 Retargeting improves upon existing scalable dynamics-aware techniques with novel components for noisy references.
首个从互联网视频到真实灵巧手部署的完整流程。 First pipeline to go from internet video to real dexterous hand rollouts.
为从业者收集人类操作数据提出有效性指南。 Proposes an efficacy playbook for practitioners collecting human data for manipulation.
在野外重建数据上达到 71%成功率,显著优于基线 25%。 Achieves 71% success rate on in-the-wild reconstructed data, significantly improving over baseline 25%.
局限 · Limitations
假设刚体物体和来自单目 RGB 的半准确度量深度。 Assumes rigid objects and semi-accurate metric depth from monocular RGB.
单目在手-物体距离上存在歧义,无法区分接触与遮挡。 Monocular ambiguity in hand-object distance, unable to distinguish contact from occlusion.
仅重建手和物体,而非完整场景,限制了对环境约束的推理。 Reconstructs only hand and object, not full scene, limiting reasoning about environmental constraints.
物理模拟器仅近似建模真实世界动力学,限制了真实世界性能的上限。 Physics simulators model real-world dynamics only approximately, bounding real-world performance.
仅 4%的互联网视频通过质量检查,表明数据损耗高。 Only 4% of internet videos pass quality check, indicating high data attrition.