In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack.
核心贡献 · Key contributions
端到端机器人学习栈,涵盖数据收集、模型设计、预训练、微调、强化学习后训练和部署。 End-to-end robot learning stack spanning data collection, model design, pre-training, SFT, RL post-training, and deployment.
定制指尖 UMI 设备与动作捕捉系统,实现亚毫米精度的人类演示。 Custom fingertip UMI device with motion-capture cage for sub-millimeter-precision human demonstrations.
具身原生混合 Transformer 骨干网络与流匹配动作专家,实现连续高频控制。 Embodied-native MoT backbone with flow-matching action expert for continuous, high-frequency control.
增量块动作表示,将策略与具身特定运动学解耦。 Delta-chunk action representation decoupling policy from embodiment-specific kinematics.
FlowPRO:无评论家、无奖励的离线强化学习,利用偏好优化与对比梯度取消。 FlowPRO: critic-free, reward-free offline RL using preference optimization with contrastive gradient cancellation.
异步推理框架与三次贝塞尔动作平滑器,实现 C1 连续部署。 Asynchronous inference framework with cubic Bézier action smoother for C1-continuous deployment.
局限 · Limitations
UMI 数据依赖动作捕捉系统,限制了野外部署的灵活性。 UMI data collected with motion-capture cage limits in-the-wild deployment flexibility.
跨具身迁移仅在有限任务和机器人上评估。 Cross-embodiment transfer evaluated only on a limited set of tasks and robots.
FlowPRO 需要遥操作干预与回滚流程来收集偏好数据。 FlowPRO requires teleoperated intervention-and-rollback pipeline for preference data collection.
部署时的执行速度与安全性权衡尚未完全解决。 Deployment-time execution speed and safety trade-offs not fully addressed.
未研究零样本泛化;数据规模可能仍不足。 Zero-shot generalization not studied; data scale may still be insufficient.