小米机器人-1:基于超过 10 万小时真实世界轨迹的视觉-语言-动作模型缩放

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

小米 MiMo 团队 Xiaomi MiMo Team · · 2026-07-16 · arXiv:2607.15330 ↗ · 被引 5

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出小米机器人-1,这是一个基础的视觉-语言-动作(VLA)模型,能够(1)遵循多样的语言指令,在未见环境中开箱即用地执行广泛的移动操作任务,以及(2)通过极少的微调数据高效适应新的下游任务。我们提出一个两阶段训练方案,包括预训练和后训练。在预训练期间,我们通过在超过 10 万小时的通过 UMI 设备收集的真实操作轨迹上训练,赋予模型广泛且可泛化的动作生成能力。关键的是,我们开发了一个可扩展的自动标注流程,用描述场景状态转换的自然语言来标注轨迹片段,为动作学习提供丰富且精确的条件。在后训练期间,我们旨在将这些能力与机器人实体和人类自然用于提示机器人的指令性指令对齐。大量实验展示了强大的缩放行为。小米机器人-1 在预训练期间随着数据规模和模型大小的增加而持续改进。这种缩放行为直接转移到后训练中,更强的预训练模型在未见环境中产生更好的开箱即用真实机器人性能。此外,小米机器人-1 作为一个强大的机器人基础策略,可以以高数据效率有效地微调复杂的、灵巧的任务。在多个模拟基准上,小米机器人-1 优于最先进的方法。值得注意的是,它在 RoboCasa365 上建立了新的最先进水平,成功率为 57.4%,超过了之前的最高 46.6%。此外,它在 RoboDojo 上获得了 20.07 的平均分,显著优于先前的最先进水平(13.07)。代码和模型检查点将发布。项目页面:https://robotics.xiaomi.com/xiaomi-robotics-1.html

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.4% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →