Hy-Embodied-0.5-VLA:从视觉-语言-动作模型到真实世界机器人学习栈

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

孙兴武 Xingwu Sun · Tencent Hunyuan · 2026-06-12 · arXiv:2606.14409 ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

在本报告中,我们介绍了 Hy-Embodied-0.5-VLA(简称 HyVLA-0.5),这是一个端到端系统,涵盖了完整的机器人学习栈:数据收集、模型设计、持续预训练和监督微调、强化学习后训练以及真实世界部署。每个组件在这个栈中都扮演着独特的角色。

In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 23)

阅读逐段中英对照全文 →