MiMo-VL 技术报告

MiMo-VL Technical Report

小米 MiMo 团队 Xiaomi MiMo Team · Xiaomi · 2025-06-04 · arXiv:2506.03569 ↗ · 被引 39

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们开源了 MiMo-VL-7B-SFT 和 MiMo-VL-7B-RL 两个强大的视觉语言模型,在通用视觉理解和多模态推理方面均达到最先进水平。MiMo-VL-7B-RL 在 40 项评估任务中的 35 项上优于 Qwen2.5-VL-7B,并在 OlympiadBench 上获得 59.4 分,超越了参数高达 78B 的模型。在 GUI 接地应用中,它以 56.1 分在 OSWorld-G 上树立了新标准,甚至优于 UI-TARS 等专用模型。我们的训练结合了四阶段预训练(2.4 万亿 token)和混合在线强化学习(MORL),集成了多样化的奖励信号。我们发现了将高质量推理数据与长思维链纳入预训练阶段的重要性,以及尽管多领域同步优化存在挑战,混合 RL 仍能带来益处。我们还贡献了一套涵盖 50 多项任务的综合评估套件,以促进可重复性和推动领域发展。模型检查点和完整评估套件可在 https://github.com/XiaomiMiMo/MiMo-VL 获取。

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of 40 evaluated tasks, and scores 59.4 on OlympiadBench, surpassing models with up to 78B parameters. For GUI grounding applications, it sets a new standard with 56.1 on OSWorld-G, even outperforming specialized models such as UI-TARS. Our training combines four-stage pre-training (2.4 trillion tokens) with Mixed On-policy Reinforcement Learning (MORL) integrating diverse reward signals. We identify the importance of incorporating high-quality reasoning data with long Chain-of-Thought into pre-training stages, and the benefits of mixed RL despite challenges in simultaneous multi-domain optimization. We also contribute a comprehensive evaluation suite covering 50+ tasks to promote reproducibility and advance the field. The model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-VL.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 19)

阅读逐段中英对照全文 →