Hy-Embodied-VLM-1.0:高效的物理世界智能体

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

孙兴武 Xingwu Sun · · 2026-07-14 · arXiv:2607.12894 ↗ · 被引 2

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

构建有能力的具身智能体不仅需要多模态感知和理解,还需要对动作进行推理、适应不断变化的情况以及与物理世界互动的智能体能力。在本报告中,我们介绍了 Hy-Embodied-VLM-1.0,这是一个高效且强大的具身基础模型,专门为在物理世界中运行的具身智能体设计。为了从预训练阶段开始培养这些能力,我们定义了一个以动作为中心的能力分类法,包含三个渐进维度:与动作相关的状态理解、动作转换推理以及顺序和自适应推理。在该分类法的指导下,我们开发了系统化的数据流程,并策划了涵盖预训练和后训练的数据组合。为了提供强大的物理世界理解和交互能力,同时支持延迟敏感部署,我们在 Hy3-A3B 语言主干和 Hy-ViT2 视觉编码器上构建了模型。其高效的混合专家架构将强大的模型容量与高推理效率相结合。我们在涵盖具身感知、物理世界理解和具身推理的 38 个基准测试的综合套件上评估了 Hy-Embodied-VLM-1.0。该模型在 38 个基准测试中的 19 个上取得了同类尺寸模型中的最佳性能,并显著优于包括 Qwen3.6-A3B 和 Cosmos 3 在内的强大竞争对手。与上一代 Hy-Embodied-0.5 MoT-2B 相比,Hy-Embodied-VLM-1.0 的平均性能提升了 8.4%。尽管仅激活了 3B 参数,但其性能接近具有 32B 激活参数的上一代模型。除了静态基准评估,Hy-Embodied-VLM-1.0 在需要多轮交互和长程推理的具身智能体任务上也表现出强大的性能。

Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce Hy-Embodied-VLM-1.0, an efficient and powerful embodied foundation model specifically designed for embodied agents operating in the physical world. To cultivate such capabilities from the pre-training stage onward, we define an action-centric capability taxonomy comprising three progressive dimensions: Action-Relevant State Understanding, Action-Transition Reasoning, and Sequential and Adaptive Reasoning. Guided by this taxonomy, we develop a systematic data pipeline and curate data mixtures spanning both pre-training and post-training. To deliver strong physical-world understanding and interaction capabilities while supporting latency-sensitive deployment, we build our model on the Hy3-A3B language backbone and the Hy-ViT2 vision encoder. Its efficient Mixture-of-Experts architecture combines strong model capacity with high inference efficiency. We evaluate Hy-Embodied-VLM-1.0 on a comprehensive suite of 38 benchmarks covering embodied perception, physical-world understanding, and embodied reasoning. The model achieves the best performance among similarly sized models on 19 of the 38 benchmarks and substantially outperforms strong competitors, including Qwen3.6-A3B and Cosmos 3. Compared with the previous-generation Hy-Embodied-0.5 MoT-2B, Hy-Embodied-VLM-1.0 improves average performance by 8.4%. Despite activating only 3B parameters, it achieves performance close to that of the previous-generation model with 32B activated parameters. Beyond static benchmark evaluation, Hy-Embodied-VLM-1.0 also demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 26)

阅读逐段中英对照全文 →