使用视频基础模型进行物理 AI 世界模拟

World Simulation with Video Foundation Models for Physical AI

范麟熙 Jim Fan · · 2025-10-28 · arXiv:2511.00062 ↗ · 被引 132

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们介绍了[Cosmos-Predict2.5],这是用于物理 AI 的 Cosmos 世界基础模型的最新版本。基于流式架构,[Cosmos-Predict2.5]在单个模型中统一了 Text2World、Image2World 和 Video2World 生成,并利用[Cosmos-Reason1](一个物理 AI 视觉语言模型)来提供更丰富的文本基础和更精细的世界模拟控制。在 2 亿个精选视频片段上训练,并通过基于强化学习的后训练进行优化,[Cosmos-Predict2.5]在视频质量和指令对齐方面相比[Cosmos-Predict1]取得了显著改进,发布了 2B 和 14B 规模的模型。这些能力使得机器人技术和自主系统能够生成更可靠的合成数据、进行策略评估和闭环模拟。我们进一步扩展了该系列,推出了[Cosmos-Transfer2.5],这是一个用于 Sim2Real 和 Real2Real 世界翻译的控制网络风格框架。尽管比[Cosmos-Transfer1]小 3.5 倍,但它提供了更高的保真度和稳健的长时域视频生成。总之,这些进展使[Cosmos-Predict2.5]和[Cosmos-Transfer2.5]成为扩展具身智能的多功能工具。为了加速物理 AI 的研究和部署,我们在 NVIDIA 开放模型许可下发布了源代码、预训练检查点和精选基准,链接为 https://github.com/nvidia-cosmos/cosmos-predict2.5 和 https://github.com/nvidia-cosmos/cosmos-transfer2.5。我们希望这些开放资源能够降低采用门槛,并促进构建下一代具身智能的创新。

We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2World, Image2World, and Video2World generation in a single model and leverages [Cosmos-Reason1], a Physical AI vision-language model, to provide richer text grounding and finer control of world simulation. Trained on 200M curated video clips and refined with reinforcement learning-based post-training, [Cosmos-Predict2.5] achieves substantial improvements over [Cosmos-Predict1] in video quality and instruction alignment, with models released at 2B and 14B scales. These capabilities enable more reliable synthetic data generation, policy evaluation, and closed-loop simulation for robotics and autonomous systems. We further extend the family with [Cosmos-Transfer2.5], a control-net style framework for Sim2Real and Real2Real world translation. Despite being 3.5$\times$ smaller than [Cosmos-Transfer1], it delivers higher fidelity and robust long-horizon video generation. Together, these advances establish [Cosmos-Predict2.5] and [Cosmos-Transfer2.5] as versatile tools for scaling embodied intelligence. To accelerate research and deployment in Physical AI, we release source code, pretrained checkpoints, and curated benchmarks under the NVIDIA Open Model License at https://github.com/nvidia-cosmos/cosmos-predict2.5 and https://github.com/nvidia-cosmos/cosmos-transfer2.5. We hope these open resources lower the barrier to adoption and foster innovation in building the next generation of embodied intelligence.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 28)

阅读逐段中英对照全文 →