Cosmos World Foundation Model Platform for Physical AI
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→物理 AI 需要先进行数字化训练。它需要自身的数字孪生(策略模型)以及世界的数字孪生(世界模型)。在本文中,我们介绍了宇宙世界基础模型平台,以帮助开发者为其物理 AI 设置构建定制化的世界模型。我们将世界基础模型定位为通用世界模型,可以微调为适用于下游应用的定制世界模型。我们的平台涵盖视频整理流水线、预训练的世界基础模型、预训练世界基础模型的后训练示例以及视频分词器。为了帮助物理 AI 构建者解决社会最关键的问题,我们将 Cosmos 开源,并通过 https://github.com/nvidia-cosmos/cosmos-predict1 提供具有宽松许可证的开放权重模型。
Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present the Cosmos World Foundation Model Platform to help developers build customized world models for their Physical AI setups. We position a world foundation model as a general-purpose world model that can be fine-tuned into customized world models for downstream applications. Our platform covers a video curation pipeline, pre-trained world foundation models, examples of post-training of pre-trained world foundation models, and video tokenizers. To help Physical AI builders solve the most critical problems of our society, we make Cosmos open-source and our models open-weight with permissive licenses available via https://github.com/nvidia-cosmos/cosmos-predict1.
物理 AI 需要首先在数字世界中进行训练。它需要自身的数字孪生——策略模型,以及世界的数字孪生——世界模型。在本文中,我们提出了 Cosmos 世界基础模型平台,以帮助开发者为他们的物理 AI 设置构建定制化的世界模型。我们将世界基础模型定位为一种通用世界模型,可以微调为适用于下游应用的定制世界模型。我们的平台涵盖视频整理流水线、预训练的世界基础模型、预训练世界基础模型的后训练示例以及视频分词器。为了帮助物理 AI 构建者解决我们社会中最关键的问题,我们将 Cosmos 开源,并通过 NVIDIA Cosmos-Predict1 提供具有宽松许可证的开放权重模型。
Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present the Cosmos World Foundation Model Platform to help developers build customized world models for their Physical AI setups. We position a world foundation model as a general-purpose world model that can be fine-tuned into customized world models for downstream applications. Our platform covers a video curation pipeline, pre-trained world foundation models, examples of post-training of pre-trained world foundation models, and video tokenizers. To help Physical AI builders solve the most critical problems of our society, we make Cosmos open-source and our models open-weight with permissive licenses available via NVIDIA Cosmos-Predict1.
物理 AI 是一种配备传感器和执行器的 AI 系统:传感器使其能够观察世界,执行器使其能够与世界交互并改变世界。它有望将人类工人从危险、费力或乏味的体力任务中解放出来。尽管近年来由于数据和算力的 Scaling(规模扩张),AI 的多个领域取得了显著进展,但物理 AI 却进展缓慢。这主要是因为为物理 AI 扩展训练数据要困难得多,因为所需数据必须包含交错观测和动作的序列。这些动作会扰动物理世界,可能对系统和世界造成严重损害。当 AI 仍处于起步阶段、探索性动作至关重要时,尤其如此。世界基础模型(WFM)——物理世界的数字孪生,物理 AI 可以安全地与之交互——一直是解决数据扩展问题的长期寻求的补救措施。
Physical AI is an AI system equipped with sensors and actuators: the sensors allow it to observe the world, and the actuators allow it to interact with and modify the world. It holds the promise of freeing human workers from physical tasks that are dangerous, laborious, or tedious. While several fields of AI have advanced significantly thanks to data and compute scaling in the recent decade, Physical AI only inches forward. This is largely because scaling training data for Physical AI is much more challenging, as the desired data must contain sequences of interleaved observations and actions. These actions perturb the physical world and may cause severe damage to the system and the world. This is especially true when the AI is still in its infancy when exploratory actions are essential. A World Foundation Model (WFM), a digital twin of the physical world that a Physical AI can safely interact with, has been a long-sought remedy to the data scaling problem.
在本文中,我们介绍了用于构建物理 AI 的 Cosmos 世界基础模型(WFM)平台。我们主要关注视觉世界基础模型,其中观测以视频形式呈现,扰动可以以各种形式存在。如图 2 所示,我们提出了一种预训练-后训练范式,将 WFM 分为预训练 WFM 和后训练 WFM。为了构建预训练 WFM,我们利用大规模视频训练数据集,使模型接触多样化的视觉体验,从而成为通才。为了构建后训练 WFM,我们使用从特定物理 AI 环境收集的数据集对预训练 WFM 进行微调,以得到针对特定物理 AI 设置的专业化 WFM。图 1 展示了我们预训练和后训练 WFM 的示例结果。
In this paper, we introduce the Cosmos World Foundation Model (WFM) Platform for building Physical AI. We are mainly concerned with the visual world foundation model, where the observations are presented as videos, and the perturbations can exist in various forms. As illustrated in fig. 2, we present a pre-training-and-then-post-training paradigm, where we divide WFMs into pre-trained and post-trained WFMs. To build a pre-trained WFM, we leverage a large-scale video training dataset to expose the model to a diverse set of visual experiences to become a generalist. To build a post-trained WFM, we fine-tune the pre-trained WFM to arrive at a specialized WFM using a dataset collected from a particular Physical AI environment for the targeted, specialized Physical AI setup. fig. 1 shows example results from our pre-trained and post-trained WFMs.
数据决定了 AI 模型的上限。为了构建高上限的预训练 WFM,我们开发了一个视频数据整理流水线。我们用它来定位具有丰富动态和高视觉质量的视频片段,这些片段有助于学习视觉内容中编码的物理规律。我们使用该流水线从 2000 万小时的视频集合中提取了约 1 亿个时长 2 到 60 秒的视频片段。对于每个片段,我们使用视觉语言模型(VLM)每 256 帧提供视频字幕。视频处理计算密集。我们利用现代 GPU 中可用的 H.264 视频编码器和解码器的硬件实现来进行解码和转码。我们的视频数据整理流水线利用了许多预训练的图像/视频理解模型。这些模型具有不同的吞吐量。为了最大化生成可训练视频数据的整体吞吐量,我们构建了一个基于 Ray 的编排流水线。细节在第 3 节中描述。
Data determines the ceiling of an AI model. To build a high-ceiling pre-trained WFM, we develop a video data curation pipeline. We use it to locate portions of videos with rich dynamics and high visual quality that facilitate learning of physics encoded in visual content. We use the pipeline to extract about 100M clips of videos ranging from 2 to 60 seconds from a 20M hour-long video collection. For each clip, we use a visual language model (VLM) to provide a video caption per 256 frames. Video processing is computationally intensive. We leverage hardware implementations of the H.264 video encoder and decoder available in modern GPUs for decoding and transcoding. Our video data curation pipeline leverages many pre-trained image/video understanding models. These models have different throughputs. To maximize the overall throughput for generating trainable video data, we build a Ray-based orchestration pipeline. The details are described in section 3.
我们探索了两种可扩展的方法来构建预训练 WFM,在第 5 节中讨论。这些方法是基于 Transformer 的扩散模型和基于 Transformer 的自回归模型。扩散模型通过逐步去除高斯噪声视频中的噪声来生成视频。自回归模型按照预设顺序,基于过去的生成逐块生成视频。这两种方法都将困难的视频生成问题分解为更简单的子问题,使其更易处理。我们利用最先进的 Transformer 架构来实现可扩展性。在第 5.1 节中,我们提出了一种基于 Transformer 的扩散模型设计,展现出强大的世界生成能力。在第 5.2 节中,我们提出了一种基于 Transformer 的自回归模型设计用于世界生成。
We explore two scalable approaches for building pre-trained WFMs discussed in section 5. These approaches are transformer-based diffusion models and transformer-based autoregressive models. A diffusion model generates videos by gradually removing noise from a Gaussian noise video. An autoregressive model generates videos piece by piece, conditioned on the past generations following a preset order. Both approaches decompose a difficult video generation problem into easier sub-problems, making it more tractable. We leverage state-of-the-art transformer architectures for their scalability. In section 5.1, we present a transformer-based diffusion model design that exhibits strong world-generation capabilities. In section 5.2, we present a transformer-based autoregressive model design for world generation.
基于 Transformer 的扩散模型和自回归模型都使用 token 作为视频的表示,前者使用向量形式的连续 token,后者使用整数形式的离散 token。我们注意到,视频的 token 化——将视频转换为一组 token 的过程——是非常不平凡的。视频包含关于视觉世界的丰富信息。然而,为了便于 WFM 的学习,我们需要将视频压缩成紧凑的 token 序列,同时最大限度地保留视频中的原始内容,因为世界基础模型训练的计算复杂度随 token 数量增长。在许多方面,构建视频 tokenizer 类似于构建视频编解码器。我们开发了一种基于注意力的编码器-解码器架构,用于学习连续和离散 token 的视频 token 化,在第 4 节中描述。
Both the transformer-based diffusion model and transformer-based autoregressive model use tokens as representations of videos, where the former uses continuous tokens in the form of vectors, and the latter uses discrete tokens in the form of integers. We note that tokenization for videos—a process that transforms videos into a set of tokens—is highly nontrivial. Video contains rich information about the visual world. However, to facilitate learning of the WFMs, we need to compress videos into sequences of compact tokens while maximally preserving the original contents in the videos as the computation complexity of world foundation model training grows with the token counts. In many ways, building a video tokenizer is similar to building a video codec. We develop an attention-based encoder-decoder architecture to learn video tokenization for both continuous and discrete tokens described in section 4.
我们在第 6 节中对预训练的世界基础模型(WFM)进行微调,以获得适用于各种 Physical AI 任务的后训练 WFM。在第 6.1 节中,我们对预训练的扩散 WFM 进行微调,使其以相机位姿为条件。这种后训练创建了一个可导航的虚拟世界,用户可以通过移动虚拟视点来探索所创建的世界。在第 6.2 节中,我们在各种机器人任务(包括视频-动作序列)上对 WFM 进行微调。我们表明,利用预训练的 WFM,我们可以根据机器人采取的动作更好地预测世界的未来状态。在第 6.3 节中,我们展示了如何将预训练的 WFM 微调用于各种自动驾驶相关任务。
We fine-tune the pre-trained WFMs to arrive at post-trained WFMs for various Physical AI tasks in Section 6. In Section 6.1, we fine-tune our pre-trained diffusion WFM to make it camera pose conditional. This post-training creates a navigable virtual world where users can explore the created world by moving the virtual viewpoint around. In Section 6.2, we fine-tune our WFMs on various robotic tasks, which consist of video-action sequences. We show that by leveraging the pre-trained WFMs, we can better predict the future state of the world based on the action taken by the robot. In Section 6.3, we demonstrate how the pre-trained WFMs can be fine-tuned for various autonomous driving-related tasks.
我们开发 WFM 的预期用途是面向 Physical AI 构建者。为了更好地保护开发者在使用世界基础模型时的安全,我们开发了一个强大的护栏系统,包括一个预护栏(pre-Guard)来阻止有害输入,以及一个后护栏(post-Guard)来阻止有害输出。详细信息在第 7 节中描述。
Our intended use of the developed WFMs is for Physical AI builders. To better protect the developers when using the world foundation models, we develop a powerful guardrail system that consists of a pre-Guard to block harmful inputs and a post-Guard to block harmful outputs. The details are described in Section 7.
我们旨在构建一个世界基础模型平台,以帮助 Physical AI 构建者推进他们的系统。为实现这一目标,我们在 NVIDIA Cosmos 上以 NVIDIA 开放模型许可协议提供我们的预训练世界基础模型和分词器。虽然本文在世界基础模型设计方面做出了若干改进,但世界基础模型问题仍远未解决。需要进一步的研究来推动最先进水平的发展。
We aim to build a world foundation model platform to help Physical AI builders advance their systems. To achieve this goal, we make our pre-trained world foundation models and tokenizers available under the NVIDIA Open Model License at NVIDIA Cosmos. While this paper makes several improvements in world foundation model design, the world foundation model problem is still far from being solved. Additional research is required to advance the state-of-the-art further.
设 \(x_{0:t}\) 为从时间 \(0\) 到 \(t\) 的真实世界视觉观测序列。设 \(c_{t}\) 为对世界的扰动。如图 3 所示,WFM 是一个模型 \(\mathcal{W}\),它基于过去的观测 \(x_{0:t}\) 和当前的扰动 \(c_{t}\) 预测未来时刻 \(t+1\) 的观测 \(\hat{x}_{t+1}\)。在我们的案例中,\(x_{0:t}\) 是 RGB 视频,而 \(c_{t}\) 是扰动,可以采取多种形式。它可以是 Physical AI 采取的动作、随机扰动、扰动的文本描述等。
Let \(x_{0:t}\) be a sequence of visual observations of the real world from time \(0\) to \(t\). Let \(c_{t}\) be the perturbation to the world. As illustrated in Fig. 3, a WFM is a model \(\mathcal{W}\) that predicts the future observation at time \(t+1\), \(\hat{x}_{t+1}\), based on the past observation \(x_{0:t}\) and the current perturbation \(c_{t}\). In our case, \(x_{0:t}\) is an RGB video, while \(c_{t}\) is a perturbation that can take many forms. It can be an action taken by the Physical AI, a random perturbation, a text description of the perturbation, etc.
我们相信,世界基础模型(WFM)在许多方面对物理 AI 构建者有用,包括(但不限于)
We believe a WFM is useful to Physical AI builders in many ways, including (but not limited to)
**策略评估**。这指的是评估物理 AI 系统中策略模型的质量。与其将训练好的策略部署到真实世界中运行的物理 AI 系统上进行评估,不如让物理 AI 系统的数字副本与世界基础模型交互。基于 WFM 的评估更具成本效益和时间效率。借助 WFM,构建者可以将策略模型部署在原本不可用的未见环境中。WFM 可以帮助开发者快速排除不合适的策略,并将物理资源集中在少数有前景的策略上。
Policy evaluation. This refers to evaluating the quality of a policy model in a Physical AI system. Instead of evaluating a trained policy by deploying it to a Physical AI system operating in the real world, one could instead let the digital copy of the Physical AI system interact with the world foundation model. The WFM-based evaluation is more cost-effective and time-efficient. With the WFM, builders can deploy the policy model in unseen environments that are otherwise unavailable. WFMs can help developers rule out incapable policies quickly and focus the physical resources on a few promising ones.
**策略初始化**。策略模型根据当前观测和给定任务生成物理 AI 系统要执行的动作。一个训练良好的 WFM,基于输入扰动建模世界的动态模式,可以作为策略模型的良好初始化。这有助于解决物理 AI 中的数据稀缺问题。
Policy initialization. A policy model generates actions to be taken by the Physical AI system based on the current observations and the given task. A well-trained WFM, which models the dynamic patterns of the world based on the input perturbations, can serve as a good initialization of the policy model. This helps address the data scarcity problem in Physical AI.
**策略训练**。WFM 与奖励模型配对,可以作为物理世界的代理,在强化学习设置中为策略模型提供反馈。智能体可以通过与 WFM 交互来获得解决任务的熟练度。
Policy training. A WFM paired with a reward model can be a proxy for the physical world to provide feedback to the policy model in a reinforcement learning setup. The agent can gain proficiency in solving tasks by interacting with the WFM.
**规划或模型预测控制**。WFM 可用于模拟物理 AI 系统采取不同动作序列后的不同未来状态。然后可以使用成本/奖励模块根据结果量化这些不同动作序列的性能。物理 AI 然后可以根据模拟结果整体执行最佳动作序列,如规划算法,或以滚动时域方式执行,如模型预测控制。世界模型的准确性决定了这些决策策略性能的上限。
Planning or model-predictive control. A WFM can be used to simulate different future states following different action sequences taken by a Physical AI system. A cost/reward module can then be used to quantify the performance of these different action sequences based on the outcomes. The Physical AI can then execute the best action sequence based on the simulation results as a whole, as in planning algorithms or in a receding horizon manner, as in model-predictive control. The accuracy of the world model upper-bounds the performance of these decision-making strategies.
合成数据生成。WFM 可用于生成用于训练的合成数据。它还可以进行微调,以条件于渲染元数据,如深度图或语义图。可以将条件 WFM 用于 Sim2Real 用例。
Synthetic data generation. A WFM can be used to generate synthetic data for training. It can also be fine-tuned to be conditioned on rendering metadata such as depth or semantic maps. One can use the conditional WFM for the Sim2Real use case.
虽然我们列出了这些可能性,但本文并未包含将 Cosmos WFM 应用于这些可能性的实证结果。我们渴望在未来的工作中验证这些主张。
While we list the possibilities, this paper does not include empirical results in applying Cosmos WFMs to them. We are eager to verify the claims in future work.
图 4 展示了本文中 Cosmos WFM 平台可用的内容,包括视频策展、视频分词化、世界基础模型预训练、世界基础模型后训练以及护栏。
Fig. 4 visualizes what is available in the Cosmos WFM platform in this paper, including video curator, video tokenization, world foundation model pre-training, world foundation model post-training, and guardrail.
视频策展。我们开发了一个可扩展的视频数据策展流水线。每个视频被分割成无场景变化的独立镜头。然后对片段应用一系列过滤步骤,以定位高质量、动态信息丰富的子集用于训练。这些高质量镜头随后使用 VLM 进行标注。我们再进行语义去重,以构建一个多样但紧凑的数据集。
Video curator. We develop a scalable video data curation pipeline. Each video is split into individual shots without scene changes. A sequence of filtering steps is then applied to the clips to locate high-quality and dynamic information-rich subsets for training. These high-quality shots are then annotated using a VLM. We then perform semantic de-duplication to construct a diverse but compact dataset.
视频分词化。我们开发了一系列不同压缩比的视频分词器。这些分词器是因果的。当前帧的令牌计算不基于未来观察。这种因果设计有几个好处。在训练方面,它使得图像和视频联合训练成为可能,因为当输入是单张图像时,因果视频分词器也是图像分词器。这对于视频模型利用图像数据集进行训练很重要,这些数据集包含丰富的世界外观信息,且往往更多样。在应用方面,因果视频分词器与生活在因果世界中的物理 AI 系统更契合。
Video tokenization. We develop a family of video tokenizers of different compression ratios. These tokenizers are causal. The token computation for the current frames is not based on future observation. This causal design has several benefits. On the training side, it makes joint image and video training possible since a causal video tokenizer is also an image tokenizer when the input is a single image. This is important for the video model to leverage image datasets for training, which contain rich appearance information of the worlds and tend to be more diverse. On the application side, causal video tokenizers are better aligned with Physical AI systems that live in the causal world.
WFM 预训练。我们探索了两种可扩展的方法来构建预训练的世界基础模型——扩散模型和自回归模型。我们使用 Transformer 架构,因为它具有可扩展性。
WFM pre-training. We explore two scalable approaches for building pre-trained world foundation models—the diffusion model and the autoregressive model. We use the transformer architecture for its scalability.
对于基于扩散的 WFM,预训练包括两个步骤:1)文本到世界生成预训练和 2)视频到世界生成预训练。具体来说,我们训练模型根据输入文本提示生成视频世界。然后我们微调它,使其根据过去的视频和输入文本提示生成未来的视频世界,我们称之为视频到世界生成任务。
For the diffusion-based WFM, the pre-training consists of two steps: 1) Text2World generation pre-training and 2) Video2World generation pre-training. Specifically, we train the model to generate a video world based on the input text prompt. We then fine-tune it to generate a future video world based on the past video and an input text prompt, which we refer to as the Video2World generation task.
对于基于自回归的 WFM,预训练包括两个步骤:1)普通的下一词预测,2)基于文本条件的 Video2World 生成。我们首先训练模型根据过去的视频输入生成未来的视频世界——即前瞻生成。然后我们对其进行微调,使其能够根据过去的视频和文本提示生成未来的视频世界。
For the autoregressive-based WFM, the pre-training consists of two steps: 1) vanilla next-token generation and 2) text-conditioned Video2World generation. We first train the model to generate a future video world based on the input of past video—foresight generation. We then fine-tune it to generate a future video world based on the past video and a text prompt.
Video2World 生成模型是一个预训练的世界模型,它根据当前观察(过去的视频)和控制输入(提示)生成未来。对于基于扩散和基于自回归的 WFM,我们构建了一系列不同能力的模型,并研究它们在各种下游应用中的有效性。
The video2world generation model is a pre-trained world model that generates the future based on the current observation (the past video) and control input (prompt). For both diffusion-based and autoregressive-based WFMs, we build a family of models with different capacities and study their effectiveness on various downstream applications.
我们进一步对预训练的扩散 WFM 进行微调,得到扩散解码器,以增强自回归模型的生成结果。为了更好地控制 WFM,我们还构建了一个基于大语言模型(LLM)的提示上采样器。
We further fine-tune our pre-trained diffusion WFM to arrive at a diffusion decoder to enhance the generation results of the autoregressive model. To better control the WFM, we also built a prompt upsampler based on a Large Language Model (LLM).
世界模型后训练。我们展示了预训练的 WFM 在多个下游物理 AI 应用中的应用。我们以相机位姿作为输入提示对预训练的 WFM 进行微调,这使我们能够在创建的世界中自由导航。我们还展示了预训练的 WFM 如何针对人形机器人和自动驾驶任务进行微调。
World model post-training. We show applications of the pre-trained WFMs on several downstream Physical AI applications. We fine-tune a pre-trained WFM with the camera pose as the input prompt. This allows us to navigate freely in the created world. We also demonstrate how our pre-trained WFMs might be fine-tuned for humanoid and autonomous driving tasks.
护栏。为了安全使用所开发的世界基础模型,我们开发了一个护栏系统,用于阻止有害的输入和输出。
Guardrail. For safe usage of the developed world foundation models, we develop a guardrail system where harmful inputs and outputs are blocked.
我们描述了我们的视频整理流程,该流程为分词器和世界基础模型(WFMs)生成高质量的训练数据集。如图 5 所示,我们的流程包括 5 个主要步骤:1)拆分,2)过滤,3)标注,4)去重,5)分片。每一步都旨在提高数据质量并满足模型训练的要求。我们首先介绍原始数据集,然后详细描述每个步骤。
We describe our video curation pipeline, which produces high-quality training datasets for both tokenizers and WFMs. As shown in Fig. 5, our pipeline consists of 5 main steps: 1) splitting, 2) filtering, 3) annotation, 4) deduplication, and 5) sharding. Every step is tailored to improve the data quality and accommodate the requirements of model training. We first present our raw dataset and then describe each step in detail.
我们使用专有视频数据集和公开可用的开放域互联网视频来训练我们的模型。我们的目标是赋能 Physical AI 开发者。为此,我们精心策划了视频训练数据集,以覆盖各种 Physical AI 应用,并针对以下视频类别:
We use both proprietary video datasets and publicly available open-domain Internet videos to train our models. Our goal is to enable Physical AI developers. To this end, we curate the video training dataset to cover various Physical AI applications and target the following video categories:
手部动作和物体操作(16%),
Hand motion and object manipulation (16%),
人体运动和活动(10%),
Human motion and activity (10%),
空间感知和导航(16%),
Spatial awareness and navigation (16%),
第一人称视角(8%),
First person point-of-view (8%),
动态相机运动(8%),
Dynamic camera movements (8%),
合成渲染(4%),以及
Synthetically rendered (4%), and
这些视频广泛覆盖了不同的视觉对象和动作。它们的多样性提升了我们世界模型(WFM)的泛化能力,并帮助模型处理各种下游任务。这些视频的非结构化特性和庞大数量,从算法和基础设施两个角度都给高效处理带来了诸多挑战。视频可能采用多种编解码器编码,并具有不同的宽高比、分辨率、长度等。许多视频还经过后处理或编辑,带有不同的视觉效果,如果处理不当,可能会在生成的视频中引入不必要的伪影,并损害世界模型的性能。
These videos offer a broad coverage of different visual objects and actions. Their diversity improves the generalization of our WFMs and helps the models handle different downstream tasks. The unstructured nature of these videos and their sheer volume creates many challenges to processing them efficiently from both an algorithmic and an infrastructural perspective. The videos can be encoded with a wide variety of codecs and have different aspect ratios, resolutions, lengths, etc. Many videos have also been post-processed or edited with different visual effects, which may induce unwanted artifacts in the generated videos and hurt the performance of the world models if not appropriately handled.
我们总共积累了约 2000 万小时的原始视频,分辨率从 720p 到 4K 不等。然而,大量视频数据要么在语义上冗余,要么不包含对学习世界物理有用的信息。因此,我们设计了一系列数据处理步骤,以从原始视频中筛选出最有价值的训练部分。我们还收集图像数据,因为联合图像和视频训练已被证明能提高生成视频的视觉质量并加速模型训练。得益于数据整理流程的模块化设计,我们可以用同一流程处理图像和视频数据,并生成用于预训练和微调的数据集。我们为预训练生成了约\(10^{8}\)个视频片段,为微调生成了约\(10^{7}\)个视频片段。
In total, we accumulate about 20M hours of raw videos with resolutions from 720p to 4k. However, a significant amount of the video data is either semantically redundant or does not contain useful information for learning the physics of the world. Hence, we design a sequence of data processing steps to find the most valuable parts of the raw videos for training. We also collect image data as joint-image-and-video training has been shown to improve the visual quality of the generated videos and accelerate the model training. Thanks to the modular design of our data curation pipeline, we can use it to process both image and video data and generate datasets for both pre-training and fine-tuning. We generate about \(10^{8}\) video clips for pre-training and about \(10^{7}\) for fine-tuning.
我们的视频长度任意,而现代深度学习模型无法直接处理超长视频。此外,许多视频包含镜头切换。它们可能从一个场景开始,然后过渡到另一个完全无关的场景,例如,从纽约市现代厨房中两人交谈切换到非洲稀树草原上狮子追逐斑马的场景。因此,根据镜头变化对每个视频进行分割,并生成视觉上一致的视频片段至关重要,这样模型才能学习到物理上合理的视觉内容转换,而非人为剪辑的转换。
Our videos have arbitrary lengths, and modern deep-learning models cannot directly consume very long videos. Also, many videos contain shot transitions. They can start from one scene and then transition to a different scene where the two scenes can be disconnected entirely, e.g., from two people talking in a modern kitchen in New York City to a scene of lions chasing zebra in an African savanna. It is important to segment each video based on its shot changes and generate visually consistent video clips so that the model can learn visual content transitions that are physically plausible instead of artificially edited.
切分旨在将任意长度的原始视频按时间分割为无镜头切换的片段。它以原始视频为输入,生成每个镜头的起始和结束帧索引。短于 2 秒的片段会被丢弃,因为它们可能是镜头过渡或视觉效果。长于 60 秒的片段会被进一步切分,以保持最大长度为 60 秒。后续的过滤步骤可以判断一个片段是否包含对学习世界物理有用的信息。
Splitting aims to temporally segment raw videos of arbitrary lengths into clips without shot changes. It takes the raw videos as input and generates each shot's start and end frame indices. Clips shorter than 2s are discarded, as they could be shot transitions or visual effects. Clips longer than 60s are further split to have a maximal length of 60s. The subsequent filtering steps can then determine whether a clip contains useful information for learning the physics of the world.
镜头边界检测是一个经典的计算机视觉问题。现有方法基于视觉特征空间的变化来检测镜头边界,但它们在如何从视频帧中学习视觉特征方面有所不同。我们在表 1 中评估了几种算法:PySceneDetect、Panda70M、TransNetV2 和 AutoShot。
Shot boundary detection is a classical computer vision problem. Existing methods detect shot boundaries based on changes in the visual feature space, but they differ in how to learn visual features from video frames. We evaluate several algorithms for the task in Table 1: PySceneDetect, Panda70M, TransNetV2, and AutoShot.
PySceneDetect 是一个流行的库,它通过对 HSV 空间中颜色直方图的时间变化进行阈值化来检测镜头变化。注意,最近的 MovieGen 工作也采用了它。Panda70M 通过基于 CLIP 嵌入的拼接和过滤增强了 PySceneDetect。而 TransNetV2 和 AutoShot 则基于神经网络,在给定 100 帧滚动输入窗口的情况下,预测每一帧为过渡帧的概率。
PySceneDetect is a popular library that detects shot changes by thresholding the temporal change of color histogram in HSV space. Note that it is also adopted by the recent MovieGen work. Panda70M augments PySceneDetect with CLIP-embedding-based stitching and filtering. TransNetV2 and AutoShot, on the other hand, are neural network-based, predicting a probability of each frame being a transition frame given a 100-frame rolling input window.
选择一种能够良好处理重度编辑视频的算法至关重要,因为这些视频通常具有复杂的镜头变化并伴有各种视觉效果。这促使我们构建一个专门的基准来评估该方法是否能从视频中生成具有干净镜头切分的片段。我们的基准(名为 ShotBench,可在 https://github.com/NVlabs/ShotBench 获取)包括现有数据集,如 RAI、BBC Planet Earth、ClipShots 和 SHOT。对于 ClipShots,我们将过渡帧定义为每个镜头注释的起点和终点的中点,以与其他数据集保持一致。
It is critical to select an algorithm that can handle heavily edited videos well, as they often have complex shot changes compounded with various visual effects. This motivates us to build a dedicated benchmark to evaluate whether the method can generate clips with clean shot cuts from videos. Our benchmark (named ShotBench, available at https://github.com/NVlabs/ShotBench) includes existing datasets, such as RAI, BBC Planet Earth, ClipShots, and SHOT. For ClipShots, we define the transition frame as the midpoint of the start and end of each shot annotation to be consistent with other datasets.
我们的视频使用了多种不同的编解码器及各种设置,这给数据整理带来了挑战。我们将镜头检测得到的每个视频片段重新编码为一致的高质量 mp4 格式,从而简化了后续的数据整理流程。统一的视频编解码器也大大提高了我们用于模型训练的数据加载器的稳定性和效率。我们采用高比特率的 h264_nvenc 编解码器,并使用快速运动和高频纹理的视频对我们的设置进行压力测试,以确保没有可感知的视觉质量下降。
Our videos use many different codecs with various settings, which poses challenges to data curation. We re-encode each video clip from shot detection into a consistent, high-quality mp4 format. This simplifies the subsequent data curation process. With a unified video codec, the stability and efficiency of our dataloader for model training are also greatly improved. We use the h264_nvenc codec with a high bitrate and stress test our setting using videos with fast motion and high-frequency texture to ensure no perceptible visual degradation.
为了最大化吞吐量,我们在表 2 中全面评估了不同的硬件和软件转码配置。现代 GPU 提供了硬件加速的视频编码和解码能力。NVIDIA L40S 同时具备解码(NVDEC)和编码(NVENC)的硬件加速器,而 NVIDIA H100 仅具备 NVDEC。为了与 L40S 在表 2 中进行公平比较,我们为 H100 配置了最大可用 CPU 核心数(28 个而非 1 个)。L40S 的吞吐量比 H100 高约 17%(0.0674 \\vs\0.0574)。在软件配置方面,从 libx264 切换到 h264_nvenc,并批量转码来自同一视频的多个片段,显著提升了吞吐量。我们观察到 ffmpeg 在充分利用 NVDEC/NVENC 加速器方面存在问题,尤其是在多 GPU 节点上。用 PyNvideoCodec 替代 ffmpeg 进行视频流转码,可以大幅提高加速器利用率,并带来最大的吞吐量提升(0.3702 \\vs\0.1026)。我们仅保留 ffmpeg 用于音频混流,而使用 PyNvideoCodec 来更好地利用 GPU 的计算能力。综合所有改进后,我们的吞吐量提升了约 6.5 倍。
We thoroughly evaluate different hardware and software configurations for transcoding to maximize the throughput in table 2. Modern GPUs provide hardware-accelerated video encoding and decoding capabilities. NVIDIA L40S has hardware accelerators for both decoding (NVDEC) and encoding (NVENC), whereas NVIDIA H100 only has NVDEC. We compensate H100 with the maximum available CPU cores (28 instead of 1) for a fair comparison with L40S in table 2. L40S has about 17% higher throughput than H100 (0.0674 \\vs\0.0574). For software configurations, switching from libx264 to h264_nvenc and transcoding multiple clips from the same video in batches significantly boost the throughput. We observe issues with ffmpeg fully utilizing NVDEC/NVENC accelerators, especially on multi-GPU nodes. Replacing ffmpeg with PyNvideoCodec for video stream transcoding leads to much higher accelerator utilization and the biggest throughput improvement (0.3702 \\vs\0.1026). We only keep ffmpeg for audio remixing and use PyNvideoCodec to better leverage the computing power in the GPUs. We achieve a \(\sim 6.5\times\) increase in throughput when combining all the improvements together.
由切分步骤生成的视频片段噪声较大,质量参差不齐,涵盖各种主题。我们设计过滤步骤以:1)移除视觉质量不达标(未达到最低要求)的视频片段;2)筛选出适合微调的高质量视频片段;3)调整数据分布以构建世界模型(WFMs)。我们通过运动过滤、视觉质量过滤、文本过滤和视频类型过滤来实现上述目标。
The video clips produced from the splitting step are noisy, with vastly different qualities covering various topics. We design the filtering step to 1) remove video clips whose visual quality fails to meet our minimal requirements, 2) select high-quality video clips suitable for fine-tuning, and 3) tailor the data distribution for building WFMs. We achieve the above goal by doing motion filtering, visual quality filtering, text filtering, and video type filtering.
我们在运动过滤中有两个主要目标:1) 移除静态或具有随机剧烈相机运动(通常来自手持相机)的视频;2) 为视频标注不同类型的相机运动(例如,平移、缩放、倾斜等),这可以提供额外信息来指导模型训练。
We have two main goals in motion filtering: 1) remove videos that are static or with random abrupt camera motion (usually from hand-held cameras) and 2) tag videos with different types of camera motion (e.g., pan, zoom, tilt, etc.), which can provide additional information to guide model training.
我们构建了一个轻量级分类器用于运动过滤。分类器的输入是从视频片段中提取的运动向量或光流序列。该分类器基于 ViT 架构,并使用带标签的视频进行训练。我们尝试了来自 h264 编解码器的运动向量、Farneback 光流算法以及 NVIDIA TensorRT 加速的光流估计网络。我们发现,基于 NVIDIA TensorRT 加速光流估计构建的分类器效果最佳,在运动过滤中实现了高分类准确率。
We build a lightweight classifier for motion filtering. The input to the classifier is a sequence of motion vectors or optical flow extracted from a video clip. The classifier is based on the ViT architecture and is trained with labeled videos. We experiment with motion vectors from h264 codec, the Farneback optical flow algorithm, and an NVIDIA TensorRT-accelerated optical flow estimation network. We find that the classifier built on top of the NVIDIA TensorRT-accelerated optical flow estimation works the best, producing high classification accuracy for motion filtering.
我们考虑两个标准:失真和外观质量,用于基于视觉质量的过滤。首先,我们移除带有失真的视频片段,例如伪影、噪声、模糊、低清晰度、过曝、欠曝等。我们使用基于 DOVER 的、在人类评分视频上训练的视频质量评估模型。该模型为每个片段给出感知质量分数,我们利用这些分数移除处于底部 15%的片段。其次,我们过滤掉外观质量低的视频片段。我们对输入片段中的采样帧应用图像美学模型。我们设定一个保守的阈值,即 3.5,因为美学对于物理 AI 来说不那么重要。
We consider two criteria, distortion and appearance quality, for visual quality-based filtering. First, we remove video clips with distortions, such as artifacts, noise, blur, low sharpness, overexposure, underexposure, etc. We use a video quality assessment model trained on human-rated videos based on DOVER. This gives a perceptual quality score per clip, and we use the scores to remove clips that are in the bottom 15%. Second, we filter out video clips with low appearance quality. We apply an image aesthetic model on sampled frames from an input clip. We set a conservative threshold, i.e., 3.5, since aesthetics are less important for Physical AI.
我们的一些视频经过后处理添加文本,以便为观众提供额外信息。我们还发现,文本往往与不同的视觉效果同时出现。我们的目标是学习世界的物理规律,因此去除带有过多此类文本的视频至关重要。请注意,我们关注的是后处理中添加的文本,而非视频拍摄场景中原本存在的文本,例如驾驶视频中的街道名称。
Some of our videos are post-processed to add text to include additional information for the viewer. We also find that text tends to co-occur with different visual effects. Our goal is to learn the physics of the world. It is crucial to remove videos with such excessive text. Note that we focus on text added in post-processing instead of text in the original scene from which the video is created, such as the street names in driving videos.
我们训练了一个基于 MLP 的二分类器来检测此类视频。分类器的输入是使用 InternVideo2 提取的视频嵌入。我们使用专有的 VLM 构建训练集,以标记正负视频。我们训练的模型在验证集上达到了较高的预测准确率。
We train an MLP-based binary classifier to detect such videos. The input to the classifier is a video embedding extracted using InternVideo2. We use a proprietary VLM to build the training set to label positive and negative videos. Our trained model achieves high prediction accuracy in the validation set.
为了调整训练数据分布并过滤掉不需要的视频类型,我们设计了一个全面的分类体系,根据内容类型和视觉风格对视频进行分类。我们训练一个分类器,为每个视频片段标注分类体系中的类别。我们通过排除可能导致生成质量差或动态不真实的特定视频类型(如抽象视觉图案、游戏画面、动画内容等)来优化数据。我们进一步通过上采样与 WFM 更相关的类别(如人类动作、人与物体交互等)和下采样不太重要的类别(如自然或风景视频)来调整数据分布。
To adjust the training data distribution and filter out unwanted video types, we design a comprehensive taxonomy that categorizes videos based on their content type and visual style. We train a classifier to label each video clip with categories from the taxonomy. We refine our data by excluding specific video types that could lead to poor generation quality or unrealistic dynamics, such as abstract visual patterns, video game footage, animated content, etc. We further adjust the data distribution by upsampling from categories that are more relevant to WFMs (e.g., human action, human and object interaction, etc.) and downsampling on categories that are less important (e.g., nature or landscape videos).
鉴于没有与我们的分类体系相匹配的现有标注数据集,我们利用专有的 VLM 为分类器创建训练和评估数据。对于每个视频片段,我们向 VLM 提供八个均匀采样的帧,并查询最合适的分类标签。利用标注数据,我们在文本过滤中使用的相同 InternVideo2 嵌入上训练一个 MLP 分类器。
Given the absence of pre-existing labeled datasets matching our taxonomy, we leverage a proprietary VLM to create training and evaluation data for the classifier. For each video clip, we prompt the VLM with eight uniformly sampled frames and query for the most appropriate taxonomy label. Using the annotated data, we train an MLP classifier on the same InternVideo2 embeddings from text filtering.
文本描述通常与图像和视频数据配对,为世界模型训练提供监督和条件。我们使用一个 VLM 为每个视频片段生成高质量且一致的描述。我们配置 VLM,使其专注于视频中的客观事实和细节。使用这种方法提供视频描述,而不是依赖 Alt 文本,也减轻了世界模型的学习负担,因为我们在训练过程中无需适应不同的文本风格或格式。
Text descriptions are usually paired with image and video data to provide supervision and conditions for world model training. We use a VLM to generate high-quality and consistent captions for each video clip. We configure the VLM in a way such that it focuses on the material facts and details in the videos. Using this approach to provide descriptions of videos instead of relying on Alt text also eases the burden of learning for world models as we do not need to adapt to different text styles or formats during training.
我们测试了几种 SOTA 方法(即 VFC、Qwen2-VL、VILA)在我们的视频上进行描述生成,基于小规模人工评估,发现 VILA 生成的描述更准确。我们使用一个内部 VILA 模型,具有 13B 参数,针对视频描述进行了微调。它拥有扩大的上下文窗口,适合处理长序列、多帧上下文,最大输入和输出 token 长度分别为 5904 和 256。为了提高推理效率,我们使用 FP8 量化的 TensorRT-LLM 引擎,与 PyTorch 半精度基线相比,吞吐量提升了 10 倍,如表 3 所示。我们提示 VILA“详细阐述视频的视觉和叙事元素”,并输入从输入片段中均匀采样的 8 帧。描述的平均长度为 559 个字符或 97 个单词。
We test several SOTA methods (i.e., VFC, Qwen2-VL, VILA) for caption generation on our videos, and find VILA generates more accurate descriptions based on a small-scale human evaluation. We use an internal VILA model with 13B parameters, fine-tuned for video captioning. It has an enlarged context window suitable for processing long, multi-frame contexts, with a max input and output token length of 5904 and 256, respectively. To improve the inference efficiency, we use an FP8-quantized TensorRT-LLM engine, resulting in a 10 \(\times\) speed-up in throughput compared to a PyTorch half-precision baseline, as shown in table 3. We prompt VILA with “Elaborate on the visual and narrative elements of the video in detail” and feed it 8 uniformly sampled frames from the input clip. The average length of captions is 559 characters or 97 words.
鉴于视频数据量庞大,训练集中可能存在重复或近似重复的样本。对数据进行去重对于构建更均衡、更多样的数据分布至关重要,同时也能提高训练效率,并减少模型记忆特定训练样本的可能性。
Given the sheer volume of our videos, there could be duplicated or near-duplicated samples in the training set. It is critical to deduplicate the data to create a more balanced and diverse data distribution. It also improves the efficiency of training and reduces the chance of memorizing specific training samples.
我们采用 SemDeDup 和 DataComp 中的方法进行可扩展的语义去重。我们复用过滤阶段计算得到的 InternVideo2 嵌入,并使用多节点 GPU 加速的 k-means 实现(\(k=10,000\))对嵌入进行聚类。我们计算每个聚类内嵌入的两两距离以识别重复项。当检测到重复视频时,我们选择分辨率最高的视频,以确保去重不会造成质量损失。为避免将整个两两距离矩阵存储在 GPU 内存中,我们按 256 个块动态计算所需的上三角矩阵和 argmax 归约。在去重过程中,我们移除了约\(30\%\)的训练数据。
We adopt the approach from SemDeDup and DataComp for scalable semantic deduplication. We reuse the InternVideo2 embeddings computed during filtering and cluster the embeddings using a multi-node GPU-accelerated implementation of k-means with \(k=10,000\). We compute the pairwise distances within each cluster of embeddings to identify duplicates. When duplicated videos are detected, we choose the video with the highest resolution to ensure no quality is lost due to deduplication. To avoid storing the entire pairwise distance matrix in GPU memory, we calculate on-the-fly the necessary upper-triangular matrix and argmax reduction in blocks of 256. We remove about \(30\%\) of training data during deduplication.
我们还利用提取的 InternVideo2 嵌入和聚类结果构建了一个视觉搜索引擎,支持以自由文本和视频查询整个训练数据集。该搜索引擎有助于调试数据问题,并理解预训练数据集与下游应用之间的差距。
We also leverage the extracted InternVideo2 embeddings and clustering results to build a visual search engine that supports querying the whole training dataset with free-form text and videos. The search engine is useful for debugging issues in our data and understanding the gap between the pre-training dataset and downstream applications.
此步骤旨在将处理后的视频片段打包成网络数据集,供我们的模型训练器直接用于训练。我们根据视频的分辨率、宽高比和长度进行分片,以与训练课程保持一致。除了预训练数据集外,我们还利用上述各种过滤器创建了质量更高的微调数据集。
This step aims to package the processed video clips into webdatasets that our model trainer can directly consume for training. We shard the videos based on their resolution, aspect ratio, and length to align with our training curriculum. Besides pre-training datasets, we also create fine-tuning datasets with even higher quality by leveraging the different filters described above.
我们的数据处理基础设施采用 AnyScale Ray 实现了一个面向地理分布式集群的流式管道系统,解决了大规模机器学习工作流中的两个关键挑战:同质节点间的高效资源利用,以及在高延迟数据源连接下的稳健运行。通过将数据传输与计算解耦,管道在远程数据存储下也能高效运行,同时内存需求随管道复杂度而非数据集大小扩展,从而支持无界流处理。
Our data processing infrastructure uses AnyScale Ray to implement a streaming pipeline system for geographically distributed clusters, addressing two key challenges in large-scale ML workflows: efficient resource utilization across homogeneous nodes and robust operation over high-latency connections to data sources. By decoupling data transfer from computation, pipelines operate efficiently with remote data storage while maintaining memory requirements that scale with pipeline complexity rather than dataset size, enabling unbounded stream processing.
我们的架构通过并行管道阶段实现互补硬件资源的并发利用,例如同时使用网络带宽进行数据摄取、NVDEC 单元进行视频解码、GPU 进行计算密集型变换。我们扩展了 Fragmentation Gradient Descent 算法以优化这种多资源分配,调度器自动缩放各个阶段,以在专用硬件加速器之间保持均衡的吞吐量。
Our architecture enables concurrent utilization of complementary hardware resources through parallel pipeline stages, for instance, simultaneously using network bandwidth for data ingestion, NVDEC units for video decoding, and GPUs for compute-intensive transformations. We extend the Fragmentation Gradient Descent algorithm to optimize this multi-resource allocation, with our scheduler automatically scaling individual stages to maintain balanced throughput across specialized hardware accelerators.
Tokenizer 是现代大规模模型的基础构建模块。它们通过以无监督方式学习一个瓶颈潜空间,将原始数据转换为更高效的表示。具体而言,视觉 tokenizer 将原始且冗余的视觉数据(如图像和视频)映射为紧凑的语义 token,这对于处理高维视觉数据至关重要。这种能力不仅使得大规模 Transformer 模型的高效训练成为可能,还使得在有限计算资源上进行推理变得更加普及。图 6 示意性地展示了 tokenization 训练流程,其目标是训练编码器和解码器,使瓶颈 token 表示能最大程度地保留输入中的视觉信息。
Tokenizers are fundamental building blocks of modern large-scale models. They transform raw data into more efficient representations by learning a bottlenecked latent space discovered in an unsupervised manner. Specifically, visual tokenizers map raw and redundant visual data—such as images and videos—into compact semantic tokens, making them crucial for handling high-dimensional visual data. This ability not only enables efficient training of large-scale transformer models but also democratizes their inference on limited computational resources. Fig. 6 schematically illustrates the tokenization training pipeline, where the goal is to train the encoder and decoder so that the bottleneck token representation maximally preserves visual information in the input.
Tokenizer 分为两种类型:连续型和离散型(见图 7 的示例)。连续型 tokenizer 将视觉数据编码为连续的潜嵌入,如潜在扩散模型(如 Stable Diffusion 或 VideoLDM)中所用。这些嵌入适用于通过从连续分布中采样来生成数据的模型。离散型 tokenizer 将视觉数据编码为离散的潜码,并将其映射为量化索引,如自回归 Transformer(如 VideoPoet)中所见。这种离散表示对于使用交叉熵损失训练的模型(如 GPT)是必需的。图 7 展示了这两种类型的 token。
Tokenizers come in two types: continuous and discrete (see Fig. 7 for illustrations). Continuous tokenizers encode visual data into continuous latent embeddings, as in latent diffusion models like Stable Diffusion or VideoLDM. These embeddings are suitable for models that generate data by sampling from continuous distributions. Discrete tokenizers encode visual data into discrete latent codes, mapping them into quantized indices, as seen in autoregressive transformers such as VideoPoet. This discrete representation is necessary for models such as GPT that are trained with the cross-entropy loss. Fig. 7 illustrates the two types of tokens.
Tokenizer 的成功在很大程度上取决于其在不损害后续视觉重建质量的前提下实现高压缩率的能力。一方面,高压缩减少了存储和计算需求。另一方面,过度压缩可能导致重要视觉细节的丢失。这种权衡给 tokenizer 设计带来了重大挑战。
The success of tokenizers largely relies on their ability to deliver high compression rates without compromising their subsequent visual reconstruction quality. On one hand, high compression reduces storage and computational demands. On the other hand, excessive compression can lead to the loss of essential visual details. This trade-off presents a significant challenge in tokenizer design.
我们提出了 Cosmos Tokenizer,这是一个视觉 tokenizer 套件,包含用于图像和视频的连续型和离散型 tokenizer。Cosmos Tokenizer 提供了卓越的视觉重建质量和推理效率。它提供多种压缩率,以适应不同的计算约束和应用需求。表 4 比较了不同的视觉 tokenizer 及其能力。
We present Cosmos Tokenizer, a suite of visual tokenizers that includes both continuous and discrete tokenizers for images and videos. Cosmos Tokenizer offers exceptional visual reconstruction quality and inference efficiency. It offers a range of compression rates to accommodate diverse computational constraints and application needs. Table 4 presents a comparison of different visual tokenizers and their capabilities.
我们采用轻量级且计算高效的架构来设计 Cosmos Tokenizer,并带有时间因果机制。具体而言,我们使用因果时间卷积层和因果时间注意力层来保持视频帧的自然时间顺序,从而确保使用单一统一网络架构对图像和视频进行无缝 tokenization。
We design Cosmos Tokenizer using a lightweight and computationally efficient architecture with a temporally causal mechanism. Specifically, we employ causal temporal convolution layers and causal temporal attention layers to preserve the natural temporal order of video frames, ensuring seamless tokenization of images and videos using a single unified network architecture.
我们直接在高质量图像和长时长视频上训练我们的 tokenizer,不限制类别或宽高比。与专注于特定数据类别和尺寸的现有 tokenizer 不同,Cosmos Tokenizer 支持多种宽高比,包括 1:1、3:4、4:3、9:16 和 16:9。它们在推理时对时间长度不敏感,能够对超过训练时时间长度的视频进行 tokenize。
We train our tokenizers directly on high-resolution images and long-duration videos without limiting the categories or aspect ratios. Unlike existing tokenizers that focus on specific data categories and sizes, the Cosmos Tokenizer operates across various aspect ratios—including 1:1, 3:4, 4:3, 9:16, and 16:9. They are temporally length-agnostic during inference, capable of tokenizing beyond the temporal length on which they were trained.
我们还在标准图像和视频基准数据集上评估了我们的 tokenizer,包括 MS-COCO 2017、ImageNet-1K 和 DAVIS。为了促进 Physical AI 应用的视频 tokenization 研究,我们整理了一个视频数据集,涵盖 Physical AI 的多种视频类别,包括鱼眼、机器人、驾驶、人类活动和空间导航。该数据集可在 github.com/NVlabs/TokenBench 获取。
We also evaluate our tokenizers on standard image and video benchmarking datasets, including MS-COCO 2017, ImageNet-1K, and DAVIS. To facilitate the video tokenization study for Physical AI applications, we curate a video dataset that covers many video categories for Physical AI, ranging from fish-eye, robotics, driving, human activities, and spatial navigation. The dataset is available at github.com/NVlabs/TokenBench.
如图 8 所示,我们的评估结果表明,Cosmos Tokenizer 显著优于现有 tokenizer,例如在 DAVIS 视频上重建质量提高了 +4 dB PSNR。它的运行速度最高可提升 \(12\times\) 倍,并且可以在单个 NVIDIA A100 GPU(80GB 内存)上一次编码长达 8 秒的 1080p 视频和 10 秒的 720p 视频,而不会耗尽内存。
As shown in Fig. 8, our evaluation results demonstrate that the Cosmos Tokenizer significantly outperforms existing tokenizers by a large margin—for instance, achieving a +4 dB PSNR improvement in reconstruction quality on DAVIS videos. It runs up to \(12\times\) faster and can encode videos up to 8 seconds at 1080p and 10 seconds at 720p in one shot without running out of memory on a single NVIDIA A100 GPU with 80GB memory.
Cosmos Tokenizer 被设计为一种编码器-解码器架构。给定输入视频 \(x_{0:T}\in\mathbb{R}^{(1+T)\times H\times W\times 3}\),其中 \(H\)、\(W\)、\(T\) 分别为高度、宽度和帧数,编码器 (\(\mathcal{E}\)) 将输入分词为 token 视频 \(z_{0:T^{\prime}}\in\mathbb{R}^{(1+T^{\prime})\times H^{\prime}\times W^{\prime}\times C}\),空间压缩因子为 \(s_{HW}=\frac{H}{H^{\prime}}=\frac{W}{W^{\prime}}\),时间压缩因子为 \(s_{T}=\frac{T}{T^{\prime}}\)。然后解码器 (\(\mathcal{D}\)) 从这些 token 重建输入视频,得到重建视频 \(\hat{x}_{0:T}\in\mathbb{R}^{(1+T)\times H\times W\times 3}\),数学上表示为:
Cosmos Tokenizer is designed as an encoder-decoder architecture. Given an input video \(x_{0:T}\in\mathbb{R}^{(1+T)\times H\times W\times 3}\), with \(H\), \(W\), \(T\) being the height, width, and number of frames, the encoder (\(\mathcal{E}\)) tokenizes the inputs into a token video \(z_{0:T^{\prime}}\in\mathbb{R}^{(1+T^{\prime})\times H^{\prime}\times W^{\prime}\times C}\), with a spatial compression factor of \(s_{HW}=\frac{H}{H^{\prime}}=\frac{W}{W^{\prime}}\) and a temporal compression factor of \(s_{T}=\frac{T}{T^{\prime}}\). The decoder (\(\mathcal{D}\)) then reconstructs the input video from these tokens, resulting in the reconstructed video \(\hat{x}_{0:T}\in\mathbb{R}^{(1+T)\times H\times W\times 3}\), mathematically given by:
我们的架构采用时间因果设计,确保每个阶段仅处理当前和过去的帧,与未来帧无关。与常见方法不同,我们的分词器在小波空间中运行,输入首先经过 2 级小波变换。具体来说,小波变换以分组方式映射输入视频 \(x_{0:T}\),在 \(x\)、\(y\) 和 \(t\) 方向上对输入进行 4 倍下采样。分组形式为:\(\{x_{0},x_{1:4},x_{5:8},...,x_{(T-3):T}\}\rightarrow\{g_{0},g_{1},g_{2},...,g_{T/4}\}\)。后续编码器阶段以时间因果方式处理帧,如 \(\{g_{0},g_{0:1},g_{0:2},...\}\rightarrow\{\xi_{0},\xi_{1},\xi_{2},...\}\)。后续编码器阶段遵循类似方案,最终输出 token \(z_{0:T^{\prime}}\)。因果设计有助于将基于分词器构建的模型适配到下游物理 AI 应用,这些应用通常在时间因果设置下运行。小波变换使我们能够在更紧凑的视频表示上操作,消除像素信息中的冗余,使其余层专注于更语义化的压缩。
Our architecture employs a temporally causal design, ensuring that each stage processes only current and past frames, independent of future frames. Unlike common approaches, our tokenizer operates in the wavelet space, where inputs are first processed by a 2-level wavelet transform. Specifically, the wavelet transform maps the input video \(x_{0:T}\) in a group-wise manner to downsample the inputs by a factor of four along \(x\), \(y\), and \(t\). The groups are formed as: \(\{x_{0},x_{1:4},x_{5:8},...,x_{(T-3):T}\}\rightarrow\{g_{0},g_{1},g_{2},...,g_{T/4}\}\). Subsequent encoder stages process the frames in a temporally causal manner as \(\{g_{0},g_{0:1},g_{0:2},...\}\rightarrow\{\xi_{0},\xi_{1},\xi_{2},...\}\). Successive encoder stages follow a similar scheme, finally outputting the tokens \(z_{0:T^{\prime}}\). The causal design helps adapt models built on top of the tokenizer to downstream Physical AI applications that often operate on the temporal causal setting. The wavelet transform allows us to operate on a more compact video representation that eliminates redundancies in pixel information, allowing the remaining layers to focus on more semantic compression.
我们的编码器阶段(小波变换后)由一系列残差块和下采样块交替实现。在每个块中,我们采用时空分解的 3D 卷积,首先应用核大小为 \(1\times k\times k\) 的 2D 卷积来捕获空间信息,然后应用核大小为 \(k\times 1\times 1\) 的时间卷积来捕获时间动态。我们使用 \(k-1\) 的左填充以确保因果性。为了捕获长距离依赖,我们利用具有全局支持区域的时空分解因果自注意力——例如,对于最后一个编码器块,支持区域为 \(1+T^{\prime}\)。我们使用 Swish 激活函数实现非线性。我们采用层归一化(LayerNorm)而非组归一化(GroupNorm),这可以防止在潜在空间或重建输出的特定区域出现大幅值。解码器与编码器镜像,将下采样块替换为上采样块。图 9 展示了 Cosmos Tokenizer 整体架构的概览。
Our encoder stages (post wavelet transform) are implemented using a series of residual blocks interleaved with downsampling blocks. In each block, we employ a spatio-temporal factorized 3D convolution, where we first apply a 2D convolution with a kernel size of \(1\times k\times k\) to capture spatial information, followed by a temporal convolution with a kernel size of \(k\times 1\times 1\) to capture temporal dynamics. We use left padding of \(k-1\) to ensure causality. To capture long-range dependencies, we utilize a spatio-temporal factorized causal self-attention with a global support region—for instance, \(1+T^{\prime}\) for the last encoder block. We use the Swish activation function for non-linearity. We leverage Layer Normalization (LayerNorm) instead of Group Normalization (GroupNorm), which prevents large magnitudes from appearing in specific regions of the latent space or reconstructed outputs. The decoder mirrors the encoder, replacing the downsampling blocks with an upsampling block. fig. 9 depicts an overview of the overall Cosmos Tokenizer architecture.
我们采用普通自编码器(AE)公式来建模连续分词器的潜在空间。对于离散分词器,我们采用有限标量量化(FSQ)作为潜在空间量化器。连续分词器的潜在维度为 \(16\),而离散分词器为 \(6\),这表示 FSQ 级别的数量,即 \((8,8,8,5,5,5)\)。此配置对应词汇表大小为 \(64{,}000\)。
We employ the vanilla autoencoder (AE) formulation to model the continuous tokenizer’s latent space. For discrete tokenizers, we adopt the Finite-Scalar-Quantization (FSQ) as the latent space quantizer. The latent dimension for the continuous tokenizers is \(16\), whereas for the discrete tokenizers, it is \(6\), which represents the number of the FSQ levels, which are \((8,8,8,5,5,5)\). This configuration corresponds to a vocabulary size of \(64{,}000\).
我们采用联合训练策略,以预设频率交替使用图像和视频的小批量数据。我们仅对分词器解码器的最终输出进行监督,不使用从潜在空间提取的辅助损失,如承诺损失或 KL 先验损失。例如,如果连续分词器采用 VAE 公式而非普通 AE,则需要 KL 先验损失;如果离散量化采用 VQ-VAE 而非 FSQ,则需要承诺损失。
We employ a joint training strategy by alternating mini-batches of images and videos at a preset frequency. We only supervise the final output of our tokenizer's decoder. We do not use auxiliary losses tapped into the latent spaces, such as commitment or KL prior losses. For example, if a VAE formulation were used for continuous tokenizers instead of the vanilla AE, one would need to have the KL prior loss. If a VQ-VAE were used for discrete quantization instead of the FSQ, one would need to have the commitment loss.
我们采用两阶段训练方案。在第一阶段,我们使用 L1 损失进行优化,该损失最小化输入与重建视频(\(\hat{x}_{0:T}\))之间的像素级 RGB 差异,公式如下:
We employ a two-stage training scheme. In the first stage, we optimize with the L1 loss that minimizes the pixel-wise RGB difference between the input and reconstructed video ( \(\hat{x}_{0:T}\) ), given by
以及基于 VGG-19 特征的感知损失,公式如下:
and the perceptual loss based on the VGG-19 features , given by,
其中\(\texttt{VGG}_{l}(\cdot)\in\mathbb{R}^{H\times W\times C}\)是预训练 VGG-19 网络第\(l\)层的特征,\(L\)是考虑的层数,\(\alpha_{l}\)是第\(l\)层的权重。
where \(\texttt{VGG}_{l}(\cdot)\in\mathbb{R}^{H\times W\times C}\) is the features from the \(l\) -th layer of a pre-trained VGG-19 network, \(L\) is the number of layers considered, and \(\alpha_{l}\) is the weight of the \(l\) -th layer.
在第二阶段,我们使用光流(OF)损失来处理重建视频的时间平滑性,
In the second stage, we use the optical flow (OF) loss to handle the temporal smoothness of reconstructed videos,
以及 Gram 矩阵(GM)损失,以增强重建图像的清晰度。
and the Gram-matrix (GM) loss to enhance the sharpness of reconstructed images,
此外,我们在微调阶段使用对抗损失,以进一步增强重建细节,尤其是在大压缩率下。
Additionally, we use adversarial loss in the fine-tuning stage to further enhance reconstruction details, particularly at large compression rates.
我们在两种压缩率下训练图像分词器(记为 CI 和 DI):\(8\times 8\) 和 \(16\times 16\)。类似地,我们在三种压缩率下训练视频分词器(记为 CV 和 DV):\(4\times 8\times 8\)、\(8\times 8\times 8\) 和 \(8\times 16\times 16\)。这里,压缩率对于图像表示为 \(H\times W\),对于视频表示为 \(T\times H\times W\),其中 \(T\) 表示时间维度,\(H\) 和 \(W\) 表示空间维度。
We train the image tokenizers (denoted as CI and DI) at two compression rates: \(8\times 8\) and \(16\times 16\). Similarly, we train the video tokenizers (denoted as CV and DV) at three compression rates: \(4\times 8\times 8\), \(8\times 8\times 8\), and \(8\times 16\times 16\). Here, the compression rates are expressed as \(H\times W\) for images and \(T\times H\times W\) for videos, where \(T\) represents the temporal dimension, and \(H\) and \(W\) represent the spatial dimensions.
对于视频分词器,我们创建了两个变体:
For the video tokenizers, we create two variants:
Cosmos-0.1-Tokenizer:使用小批量采样较少数量的 720p 视频帧进行训练(CV 为 49 帧,DV 为 17 帧)。
Cosmos-0.1-Tokenizer: Trained using mini-batches sampling a smaller number of 720p video frames (49 frames for CV and 17 frames for DV).
Cosmos-Tokenize1:使用小批量采样大量 360p 或 720p 视频帧进行训练(CV 为 121 帧,DV 为 49 帧)。
Cosmos-Tokenize1: Trained using mini-batches sampling a larger number of 360p or 720p video frames (121 frames for CV and 49 frames for DV).
这种方法确保了处理图像和视频数据时在不同时间和空间分辨率上的灵活性。我们的实验表明,分词器能够很好地泛化到其训练过的分辨率,并在更高分辨率下保持高质量。
This approach ensures flexibility in handling varying temporal and spatial resolutions for image and video data. Our experiments indicate that the tokenizers generalize well to resolutions they were trained on and maintain strong quality at higher resolutions.
我们在各种图像和视频基准数据集上广泛评估了我们的 Cosmos Tokenizer 套件。对于图像分词器的评估,我们遵循先前的工作,在 MS-COCO 2017 和 ImageNet-1K 上进行评估。我们使用 MS-COCO 2017 验证子集的 5,000 张图像和 ImageNet-1K 验证子集的 50,000 张图像作为图像评估基准。
We extensively evaluate our Cosmos Tokenizer suite on various image and video benchmark datasets. For the evaluation of image tokenizers, we follow prior art to evaluate MS-COCO 2017 and ImageNet-1K. We use the MS-COCO 2017 validation subset of 5,000 images, and ImageNet-1K validation subset of 50,000 images as image evaluation benchmark.
TokenBench。对于视频分词器的评估,目前还没有针对高分辨率和长时长视频的标准基准。为此,我们引入了一个名为 TokenBench 的基准,以覆盖包括机器人操作、驾驶、第一人称和网络视频在内的广泛领域,并标准化评估。我们借助了常用于各种任务的现有视频数据集,包括 BDD100K、EgoExo-4D、BridgeData V2 和 Panda-70M。我们从每个数据集中随机采样 100 个视频,并通过取前 10 秒并将短边调整为 1080 进行预处理。对于 Panda-70M,我们手动过滤掉低质量内容和运动较小的视频。对于 EgoExo-4D,我们随机选择 100 个场景,并分别采样一个第一人称视频和一个第三人称视频。这样总共得到 500 个视频。TokenBench 的一些示例可在图 10 中找到。我们在 github.com/NVlabs/TokenBench 发布了 TokenBench。
TokenBench. For video tokenizer evaluation, there is not yet a standard benchmark for high-resolution and long-duration videos. To this end, we introduce a benchmark called TokenBench to cover a wide variety of domains, including robotic manipulation, driving, egocentric, and web videos, and standardize the evaluation. We resort to existing video datasets that are commonly used for various tasks, including BDD100K, EgoExo-4D, BridgeData V2, and Panda-70M. We randomly sample 100 videos from each dataset and preprocess them by taking the first 10 seconds and resizing the short size to 1080. For Panda-70M, we manually filter out the videos with low-quality content and small motions. For EgoExo-4D, we randomly pick 100 scenes and sample one egocentric video and one exocentric video. This results in a total of 500 videos. Some examples of TokenBench can be found in fig. 10. We release TokenBench at the github.com/NVlabs/TokenBench.
除了 TokenBench,我们还在 DAVIS 数据集上以 1080p 分辨率评估了我们的视频分词器。
In addition to TokenBench, we also evaluate our video tokenizers on the DAVIS dataset at 1080p resolution.
基线和评估指标。我们在不同的压缩率下评估我们的分词器,以展示其在不同计算需求下的有效性。我们将每个分词器与最先进的图像和视频分词器进行比较。表 4 列出了我们在各种设置下比较的具体 SOTA 分词器。评估指标包括峰值信噪比(PSNR)、结构相似性(SSIM)、图像的重建 Fréchet Inception 距离(rFID)和视频的重建 Fréchet Video 距离(rFVD)。
Baselines and evaluation metrics. We evaluate our tokenizers at various compression rates to showcase their effectiveness for different computational needs. We compare each of these tokenizers with state-of-the-art image and video tokenizers. table 4 presents the specific SOTA tokenizers we compared against in various settings. The evaluation metrics include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), reconstruction Fréchet Inception Distance (rFID) for images, and reconstruction Fréchet Video Distance (rFVD) for videos.
定量结果。表 6 和表 6 总结了连续和离散视频分词器在各种基准上的平均定量指标。如两个表所示,在 DAVIS 视频数据集和 TokenBench 上,Cosmos Tokenizer 在所有指标上均达到了最先进的性能,空间-时间压缩比为 4×8×8。此外,即使压缩率高出 2 倍和 8 倍(即 8×8×8 和 8×16×16),Cosmos Tokenizer 仍然比先前的方法取得更好的质量,展示了出色的压缩-质量权衡。
Quantitative results. tables 6 and 6 summarize the average quantitative metrics of continuous and discrete video tokenizers on various benchmarks. As shown in both tables, Cosmos Tokenizer achieves state-of-the-art performance in all the metrics compared to prior arts on both the DAVIS video dataset and TokenBench, with a spatial-temporal compression ratio of 4×8×8. Moreover, even with 2× and 8× higher compression ratios (i.e., 8×8×8 and 8×16×16), Cosmos Tokenizer still achieves better quality than prior art, showcasing an excellent compression-quality trade-off.
在多种图像和视频基准数据集上的定量结果证实,Cosmos Tokenizer 能够以较大的时空压缩比更好地表示视觉内容。
These quantitative results on a variety of image and video benchmark datasets confirm that Cosmos Tokenizer is able to better represent visual content with large spatial-temporal compression.
运行性能。表 9 展示了在单个 A100 80GB GPU 上测得的参数量以及每张图像或每帧视频的平均编码和解码时间。同时,我们还列出了先前最先进 tokenizer 的参数量和平均速度。如图所示,对于图像和视频 tokenizer,Cosmos Tokenizer 在保持最小模型尺寸的同时,速度比先前方法快 \(2\times\sim 12\times\),表明 Cosmos Tokenizer 在编码和解码视觉内容方面具有高效率。
Runtime performance. Table 9 shows the number of parameters and the averaged encoding and decoding times per image or per video frame, measured on a single A100 80GB GPU. In comparison, we also list the parameters and the average speeds of prior state-of-the-art tokenizers. As shown, for both image and video tokenizers, Cosmos Tokenizer is \(2\times\sim 12\times\) faster while maintaining the smallest model size compared to prior arts, showing that Cosmos Tokenizer has high efficiency for encoding and decoding visual content.
预训练的世界基础模型(WFM)是通才,能够捕捉真实世界物理和自然行为的通用知识。我们利用两种不同的可扩展深度学习范式——扩散模型和自回归模型——来构建两个 WFM 系列。扩散模型和自回归模型都将困难的生成问题分解为一系列更简单的子问题,并极大地推动了生成模型的发展。对于扩散模型,困难的生成问题被划分为一系列去噪问题;对于自回归模型,则被划分为一系列下一个词预测问题。我们讨论了在构建预训练 WFM 的过程中,如何利用针对现代 GPU 定制的各种并行化技术来扩展这些深度学习范式。我们使用一个由 10,000 块 NVIDIA H100 GPU 组成的集群,在三个月的时间内训练了本文报告的所有 WFM 模型。
Pre-trained WFMs are generalists that capture general knowledge of real-world physics and natural behaviors. We exploit two different scalable deep learning paradigms, diffusion models and autoregressive models, to build two families of WFMs. Both diffusion models and autoregressive models break a difficult generation problem into a sequence of easier sub-problems and have been turbo-charging the development of generative models. In the case of diffusion models, the difficult generation problem is divided into a sequence of denoising problems. In the case of autoregressive models, the difficult generation problem is divided into a sequence of next-token prediction problems. We discuss how we scale these deep learning paradigms using various parallelization techniques tailored for modern GPUs in our endeavor of building pre-trained WFMs. We train all of the WFM models reported in the paper using a cluster of 10,000 NVIDIA H100 GPUs in a time span of three months.
在表 10 中,我们展示了预训练 WFM 及其配套模型的映射。对于基于扩散的 WFM 系列,我们首先分别构建了 7B 和 14B 的 Text2World 模型,即 Cosmos-Predict1-7B-Text2World 和 Cosmos-Predict1-14B-Text2World。这些模型能够将文本提示映射到视觉世界的视频。然后,我们对 Text2World 模型进行微调,使其能够接受额外的视频输入,以表示当前观测。结果是 Video2World 模型,它根据当前观测(输入视频)和扰动(文本提示)预测未来视频。这些扩散模型是潜在扩散模型,采用连续令牌。我们使用 Cosmos-Tokenize1-CV8 × 8 × 8-720p 来生成视觉令牌。WFM 的训练文本提示由 VLM 通过视频描述生成产生。这些描述与人类对视频的描述分布不同。为了缓解领域差距,我们基于 Mistral-NeMo-12B-Instruct 模型构建了 Cosmos-UpsamplePrompt1-12B-Text2World,以帮助将人类文本提示转换为我们的扩散 WFM 偏好的提示。
In Table 10, we present a map of our pre-trained WFMs and their companions. For the diffusion-based WFM family, we start by building two Text2World models of 7B and 14B, respectively, which render Cosmos-Predict1-7B-Text2World and Cosmos-Predict1-14B-Text2World. These models can map text prompts to videos of visual worlds. We then fine-tune the Text2World models to take additional video input, representing the current observation. The result is a Video2World model where the future video is predicted based on the current observation (input video) and the perturbation (text prompt). These diffusion models are latent diffusion models that take continuous tokens. We use Cosmos-Tokenize1-CV8 × 8 × 8-720p to produce the visual tokens. The training text prompts for the WFMs are produced by a VLM through video description generation. These descriptions follow a different distribution of human descriptions of videos. To mitigate the domain gap, we build Cosmos-UpsamplePrompt1-12B-Text2World based on the Mistral-NeMo-12B-Instruct model to help convert human text prompts to those preferred by our diffusion-based WFMs.
对于基于自回归的 WFM 系列,我们首先分别构建了 4B 和 12B 的两个基础模型,仅基于当前视频观测来预测未来视频。我们分别将其命名为 Cosmos-Predict1-4B 和 Cosmos-Predict1-12B。这些是 Llama3 风格的 GPT 模型,从头开始训练用于视频预测任务,不具备语言理解能力。为了使基于自回归的 WFM 能够利用文本信息进行下一个词预测,我们通过添加到 Transformer 块中的交叉注意力层,将输入文本提示的 T5 嵌入融入 WFM。这些自回归 WFM 使用 Cosmos-Tokenize1-DV8 × 16 × 16-720p,将输入视频映射为少量整数。分词器的重度压缩有时会导致不期望的失真。为了解决这个问题,我们通过微调 Cosmos-Predict1-7B-Text2World 模型构建了一个扩散解码器(Cosmos-Predict1-7B-Decoder-DV8 × 16 × 16ToCV8 × 8 × 8-720p),将 DV8 × 16 × 16 空间中的离散令牌映射到 CV8 × 8 × 8 空间中的连续令牌。
For the autoregressive-based WFM family, we first build two base models that are 4B and 12B in size, respectively, to predict future videos purely based on the current video observation. We name them Cosmos-Predict1-4B and Cosmos-Predict1-12B, respectively. These are Llama3-style GPT models trained from scratch for the video prediction task and bear no language understanding. To enable autoregressive-based WFMs to utilize textual information for next token prediction, we incorporate T5 embeddings of the input text prompt into the WFMs through cross-attention layers added to the transformer blocks. These autoregressive WFMs use Cosmos-Tokenize1-DV8 × 16 × 16-720p, which maps an input video to a few integers. The heavy compression of the tokenizer can sometimes lead to undesired distortions. To address the problem, we build a diffusion decoder (Cosmos-Predict1-7B-Decoder-DV8 × 16 × 16ToCV8 × 8 × 8-720p) through fine-tuning the Cosmos-Predict1-7B-Text2World model to map discrete tokens in the DV8 × 16 × 16 space to continuous tokens in the CV8 × 8 × 8 space.
我们的基于扩散的世界基础模型(WFM)是潜在扩散模型,它们在分词器学习到的潜在空间中运行,从而实现对视频的紧凑、降维表示。这种设计选择具有多个优势:它降低了训练和推理期间的计算成本,同时简化了去噪任务。为了将视频分词为潜在表示,我们采用了 Cosmos-Tokenize1-CV8 \times 8 \times 8-720p。
Our diffusion-based WFMs are latent diffusion models that operate within a learned latent space of a tokenizer, enabling a compact, reduced-dimensional representation of videos. This design choice offers several advantages: it reduces computational costs during both training and inference while simplifying the denoising task. To tokenize videos into latent representations, we employ Cosmos-Tokenize1-CV8 \times 8 \times 8-720p.
为了训练我们的扩散世界基础模型(WFM),我们采用了 EDM 中提出的方法。去噪器 \(D_{\theta}\) 在噪声水平 \(\sigma\) 下的去噪分数匹配损失定义为:
To train our diffusion WFMs, we adopt the approach outlined in EDM. The denoising score matching loss for the denoiser \(D_{\theta}\), evaluated at a noise level \(\sigma\), is defined as
其中 \({\mathbf{x}}_{0}\sim p_{\rm{data}}\) 是从训练集中采样的干净图像或视频,\({\mathbf{n}}\sim{\mathcal{N}}\big(\mathbf{0},\sigma^{2}{\mathbf{I}}\big)\) 是独立同分布的高斯噪声,\(D_{\theta}\) 是一个以噪声为条件的神经网络,负责对损坏样本 \({\mathbf{x}}_{0}+{\mathbf{n}}\) 进行去噪。我们遵循 EDM 中引入的预处理设计来参数化 \(D_{\theta}\)。整体训练损失定义为 \({\mathcal{L}}(D_{\theta};\sigma)\) 在噪声水平上的加权期望:
where \({\mathbf{x}}_{0}\sim p_{\rm{data}}\) is a clean image or video sampled from the training set, \({\mathbf{n}}\sim{\mathcal{N}}\big(\mathbf{0},\sigma^{2}{\mathbf{I}}\big)\) is i.i.d. Gaussian noise, and \(D_{\theta}\) is a noise-conditioned neural network tasked with denoising the corrupted sample \({\mathbf{x}}_{0}+{\mathbf{n}}\). We adhere to the preconditioning design introduced in EDM for parameterizing \(D_{\theta}\). The overall training loss is defined as a weighted expectation of \({\mathcal{L}}(D_{\theta};\sigma)\) over the noise levels:
其中噪声水平 \(\sigma\) 的分布由超参数 \(P_{\text{mean}}\) 和 \(P_{\text{std}}\) 控制。\(\sigma_{\text{data}}\) 是训练数据的标准差,加权函数 \(\lambda(\sigma)\) 确保在训练开始时每个噪声水平的贡献相等。然而,随着训练的进行,这种平衡可能会恶化。为了缓解这个问题,我们将不同噪声水平上的优化视为一种多任务学习。我们引入 \(u(\sigma)\) 作为连续的不确定性函数,量化噪声水平 \(\sigma\) 下去噪目标 \({\mathcal{L}}(D_{\theta},\sigma)\) 的不确定性,并采用基于不确定性的加权方法。我们使用一个简单的 MLP 来参数化 \(u(\sigma)\),并在训练过程中最小化整体损失 \({\mathcal{L}}(D_{\theta})\)。直观地说,如果模型对任务不确定,即 \(u(\sigma)\) 较高,则噪声水平 \(\sigma\) 的损失贡献会被降低权重。同时,模型会因这种不确定性而受到惩罚,鼓励 \(u(\sigma)\) 尽可能低。
where the distribution of noise levels \(\sigma\) is controlled by hyperparameters \(P_{\text{mean}}\) and \(P_{\text{std}}\). \(\sigma_{\text{data}}\) is the standard deviation of the training data, and the weighting function \(\lambda(\sigma)\) ensures equal contribution of each noise level at the beginning of the training. However, as training progresses, this balance may deteriorate. To mitigate this issue, we treat the optimization over various noise levels as a form of multi-task learning. We utilize the uncertainty-based weighting approach by introducing \(u(\sigma)\) as a continuous uncertainty function quantifying the uncertainty for the denoising objective \({\mathcal{L}}(D_{\theta},\sigma)\) at noise level \(\sigma\). We use a simple MLP to parameterize \(u(\sigma)\) and minimize the overall loss \({\mathcal{L}}(D_{\theta})\) during training. Intuitively, the contribution of loss at noise level \(\sigma\) is weighted down if the model is uncertain about the task, i.e., if \(u(\sigma)\) is high. At the same time, the model is penalized for this uncertainty, encouraging \(u(\sigma)\) to be as low as possible.
与采用高斯流匹配公式的最新视频生成模型相比,我们的工作源自扩散分数匹配的视角。然而,如所示,这些框架在理论上是等价的,在目标和训练过程上具有根本的相似性。我们基于 EDM 的公式与这些见解一致,主要区别在于预处理设计和超参数的选择。在实践中,我们没有遇到 EDM 公式的任何性能限制。
Compared to recent video generative models that adopt the Gaussian flow matching formulation, our work is derived from the diffusion score matching perspective. However, as shown by, these frameworks are theoretically equivalent, sharing fundamental similarities in their objectives and training procedures. Our EDM-based formulation aligns with these insights, mainly differing in the choice of preconditioning designs and hyperparameters. In practice, we have not encountered any performance limitations with the EDM formulation.
在本节中,我们描述我们的去噪网络 \(D_{\theta}\) 的设计,该网络基于 DiT,而 DiT 最初是为标签条件图像生成而设计的。我们调整其架构以更好地适应可控视频生成的目标。我们在图 11 中展示了整体网络设计。
In this section, we describe the design of our denoiser network \(D_{\theta}\) that builds upon DiT, which was originally designed for label-conditioned image generation. We adapt its architecture to better suit our goal of controllable video generation. We visualize the overall network design in Fig. 11.
3D 分块化。我们网络的输入是形状为 \(T\times C\times H\times W\) 的潜在表示,适用于图像和视频数据,其中图像被视为单帧视频。为了准备去噪网络的输入,我们首先使用线性层对状态进行“分块化”,然后将其展平。该过程涉及将形状为 \((p_{t},p_{h},p_{w})\) 的非重叠立方体投影为网络的单个 token 输入。因此,在分块化之后,图像或视频被重塑为一维时空序列,长度为 \(THW/(p_{t}p_{h}p_{w})\)。对于我们的去噪网络,我们使用 \(p_{t}=1,p_{h}=p_{w}=2\)。
3D patchification. The input to our network is a latent representation of shape \(T\times C\times H\times W\) for both image and video data, with images differentiated by a video with a single frame. To prepare inputs for our denoiser network, we first “patchify” the state using a linear layer and subsequently flatten it. This process involves projecting non-overlapping cubes of shape \((p_{t},p_{h},p_{w})\) into individual token inputs for the network. Consequently, after patchification, an image or video is reshaped into a one-dimensional, spatiotemporal sequence of length \(THW/(p_{t}p_{h}p_{w})\) . We use \(p_{t}=1,p_{h}=p_{w}=2\) for our denoiser network.
混合位置嵌入:FPS 感知的 3D RoPE 与可学习嵌入。我们采用 3D 分解的旋转位置嵌入(RoPE)以支持任意大小、宽高比和视频长度的生成。具体来说,我们将特征维度划分为三个近似相等的部分,分别沿时间、高度和宽度轴应用 RoPE 的位置信息。在实践中,这可以通过在各自轴上拼接频率嵌入并重用为大型语言模型(LLM)优化的 RoPE 内核来高效实现,而无需在每个块中进行拆分和拼接。为了进一步支持不同帧率的视频合成,我们根据训练视频的每秒帧数(FPS)重新缩放时间频率。由于 RoPE 的相对位置编码特性和我们的 3D 分解设计,FPS 感知设计与我们联合图像-视频训练兼容。RoPE 的另一个好处在渐进训练中改变分辨率或视频长度时体现出来。通过利用神经正切核(NTK)-RoPE,我们观察到模型快速收敛,即使在 \(5{,}000\) 训练步内也能达到合理的性能。此外,我们发现每个 Transformer 块添加额外的可学习绝对位置嵌入可以进一步增强模型,降低训练损失,并减少生成视频中的变形伪影。
Hybrid positional embedding with FPS-aware 3D RoPE and learnable embedding. We employ a 3D-factorized Rotary Position Embedding (RoPE) to allow the generation of arbitrary size, aspect ratio, and video length. Specifically, we partition the feature dimension into three approximately equal chunks, each applying RoPE with positional information along the temporal, height, and width axes, respectively. In practice, this can be implemented efficiently without splitting and concatenation in each block by concatenating frequency embeddings in their respective axes and reusing RoPE kernels optimized for Large Language Models (LLMs). To further support video synthesis with varying frame rates, we rescale temporal frequencies based on the training video’s Frames Per Second (FPS). Due to RoPE’s relative positional encoding property and our 3D factorization design, the FPS-aware design is compatible with our joint image-video training. An additional benefit of RoPE is evident during progressive training when we alter resolution or video length. By leveraging Neural Tangent Kernel (NTK)-RoPE, we observe rapid model convergence, achieving reasonable performance even within \(5{,}000\) training steps. Additionally, we find that adding an extra learnable absolute positional embedding per transformer block can further enhance the model, reduce training loss, and reduce morphing artifacts in generated videos.
用于文本条件的交叉注意力。我们的网络依赖交叉注意力层来整合语言信息。每个 Transformer 块由顺序的自注意力、交叉注意力和前馈层组成。自注意力作用于时空 token,而交叉注意力使用 T5-XXL 嵌入作为键和值来整合语义上下文,从而实现有效的文本条件控制。
Cross-attention for text conditioning. We rely on cross-attention layers in our network for incorporating linguistic information. Each transformer block consists of sequential self-attention, cross-attention, and feed-forward layers. While self-attention operates over spatiotemporal tokens, cross-attention integrates semantic context using T5-XXL embeddings as keys and values, enabling effective text conditioning.
查询-键归一化。在训练的早期阶段,我们观察到注意力 logits 增长不稳定,导致注意力熵崩溃。我们遵循现有文献,在注意力操作之前对查询 \(Q\) 和键 \(K\) 进行归一化。我们在网络中的所有自注意力和交叉注意力层中使用具有可学习尺度的均方根归一化(RMSNorm)。
Query-key normalization. In the early stages of training, we observe instability in the growth of attention logits, leading to a collapse of attention entropy. We follow existing literature to normalize query \(Q\) and key \(K\) before the attention operation. We use Root Mean Square Normalization (RMSNorm) with learnable scales for all self-attention and cross-attention layers within our network.
AdaLN-LoRA。我们发现,DiT 的自适应层归一化(AdaLN)层占据了模型参数的很大一部分,但在 FLOPs 方面的计算复杂度贡献却微乎其微。受 W.A.L.T 启发,我们实现了低秩适应(LoRA),将这些层中的密集线性投影分解为低秩近似。对于 Cosmos-Predict1-7B,这种架构优化实现了参数数量减少 36%(从 11B 降至 7B 参数),同时在所有评估指标上保持性能一致,证明了我们参数高效设计的有效性。
AdaLN-LoRA. We find that DiT's adaptive layer normalization (AdaLN) layers account for a significant portion of the model parameters while contributing negligibly to the computational complexity in terms of FLOPs. Inspired by W.A.L.T, we implement Low-Rank Adaptation (LoRA) to decompose the dense linear projections in these layers into low-rank approximations. For Cosmos-Predict1-7B, this architectural optimization achieves a 36% reduction in parameter count (from 11B to 7B parameters) while maintaining performance parity across all evaluation metrics, demonstrating the effectiveness of our parameter-efficient design.
本节概述了我们在涵盖多种模态、分辨率、宽高比和条件输入的数据集上训练模型所采用的方法。
This section outlines the methodologies employed to train our models on datasets spanning multiple modalities, resolutions, aspect ratios, and conditioning inputs.
图像与视频联合训练。为了在模型训练中充分利用海量高质量、多样化的图像数据集,我们实现了一种交替优化策略,交错使用图像和视频数据批次。为了促进图像与视频领域之间的跨模态知识迁移,我们采用了一种特定于领域的归一化方案,利用分别针对图像和视频数据估计的充分统计量来对齐潜在分布。这一方法的动机源于观察到减少图像和视频潜在表示之间的分布偏移能够提升生成质量。此外,我们观察到视频潜在表示在时间和通道维度上存在非平稳统计特性。为解决这种异质性,我们采用了一种归一化策略,对视频潜在表示进行逐帧和逐通道的标准化,有效促使其更接近各向同性的高斯先验分布。
Joint image and video training. To leverage the vast abundance of high-quality, diverse image datasets in model training, we implement an alternating optimization strategy that interleaves batches of image and video data. To facilitate cross-modal knowledge transfer between image and video domains, we adopt a domain-specific normalization scheme that aligns the latent distributions using sufficient statistics estimated independently for image and video data. This approach is motivated by the observation that reducing the distributional shift between image and video latent representations improves generation quality. Furthermore, we observe non-stationary statistics across temporal and channel dimensions in video latent representations. To address this heterogeneity, we employ a normalization strategy that applies frame-wise and channel-wise standardization to video latent representations, effectively encouraging them to better approximate an isotropic Gaussian prior distribution.
除了跨模态知识迁移,我们的归一化方案还提供了一个重要的理论优势:训练过程中信噪比的尺度不变性。考虑两个均值为零但尺度不同的潜在表示:一个标准化为单位方差,另一个方差为 4。当添加高斯噪声 \({\mathcal{N}}(0,\sigma^{2})\) 以达到标准化表示所需的信噪比时,对于未归一化的表示,我们必须将噪声缩放为 \({\mathcal{N}}(0,4\sigma^{2})\) 以保持相同的比率。通过对所有潜在表示进行标准化,我们确保了不同尺度下信噪比的一致性,从而即使在训练过程中更新底层分词器,也能促进模型的适应。
Beyond cross-modality knowledge transfer, our normalization scheme provides an important theoretical benefit: scale invariance in the signal-to-noise ratio during training. Consider two zero-mean latent representations with different scales: one standardized to unit variance, and another with variance 4. When adding Gaussian noise \({\mathcal{N}}(0,\sigma^{2})\) to achieve a desired signal-to-noise ratio for the standardized representation, we must scale the noise to \({\mathcal{N}}(0,4\sigma^{2})\) for the unnormalized representation to maintain the same ratio. By standardizing all latent representations, we ensure consistent signal-to-noise ratios across different scales, facilitating model adaptation even when the underlying tokenizer is updated during training.
为了保持计算效率,我们平衡图像和视频批次的大小,以确保 GPU 上的内存利用率相当。然而,我们观察到视频批次的去噪损失收敛速度比图像批次损失慢。我们将此归因于视频帧固有的时间冗余,这导致视频批次的梯度幅度较小。借鉴多分辨率图像训练的最新进展,我们通过将视频批次的噪声水平按帧数的平方根相对于图像批次噪声水平进行缩放,来解决这种收敛差异。
To maintain computational efficiency, we balance image and video batch sizes to ensure comparable memory utilization across GPUs. However, we observe that the video batch denoising loss exhibits slower convergence compared to the image batch loss. We attribute this to the inherent temporal redundancy in video frames, which results in smaller gradient magnitudes for video batches. Drawing inspiration from recent advances in multi-resolution image training , we address this convergence discrepancy by scaling the video batch noise levels by the square root of the frame count relative to image batch noise levels.
\(10{,}240\)(上下文长度)的计算方式为:\(640\)(宽度)\(\div 8\)(分词)\(\div 2\)(分块)\(\times 512\)(高度)\(\div 8\)(分词)\(\div 2\)(分块)\(\times[(57-1)\div 8+1]\)(帧分词)。
\(10{,}240\) (the context length) is computed as: \(640\) (width) \(\div 8\) (tokenize) \(\div 2\) (patchify) \(\times 512\) (height) \(\div 8\) (tokenize) \(\div 2\) (patchify) \(\times[(57-1)\div 8+1]\) (tokenize frames).
\(56{,}320\)(上下文长度)的计算方式为:\(1280\)(宽度)\(\div 8\)(分词)\(\div 2\)(分块)\(\times 704\)(高度)\(\div 8\)(分词)\(\div 2\)(分块)\(\times[(121-1)\div 8+1]\)(帧分词)。
\(56{,}320\) (the context length) is computed as: \(1280\) (width) \(\div 8\) (tokenize) \(\div 2\) (patchify) \(\times 704\) (height) \(\div 8\) (tokenize) \(\div 2\) (patchify) \(\times[(121-1)\div 8+1]\) (tokenize frames).
渐进式训练。我们采用渐进式训练策略,各阶段的具体细节见表 12。初始阶段在 512 像素分辨率下对视频和图像进行训练,视频由 57 帧组成。随后,我们过渡到 720 像素的目标分辨率,将视频长度增加到 121 帧。在海量数据上进行预训练后,我们在高质量子集上对模型进行微调,迭代次数为\(\mathcal{O}(10k)\),学习率线性衰减。与文献[引用]的发现一致,我们也发现微调可以提高生成视频的质量。
Progressive training. We adopt a progressive training strategy, with the specifics of each stage detailed in Table 12. The initial stage involves training on videos and images at a resolution of 512 pixels, using videos composed of 57 frames. Subsequently, we transition to the target resolution of 720 pixels, increasing the video length to 121 frames. After pre-training on massive data, we fine-tune the model on a high-quality subset for \(\mathcal{O}(10k)\) iterations with a linearly decaying learning rate. Consistent with findings from [reference], we also find that fine-tuning can improve the quality of the generated videos.
多宽高比训练。为了适应不同宽高比的内容,我们将数据组织为五个不同的桶,对应 1:1、3:4、4:3、9:16 和 16:9 的比例,并将每个图像或视频分配到最接近其宽高比的桶中。训练时,每个数据并行进程组从一个桶中采样,允许不同并行进程组使用不同的桶。我们采用最长边缩放,以最大程度保留提示中描述的原始内容信息。对于批处理,我们对缺失像素应用反射填充,并将填充掩码提供给扩散主干,从而在推理时实现精确控制。
Multi-aspect training. To accommodate content with varying aspect ratios, we organize the data into five distinct buckets corresponding to ratios of 1:1, 3:4, 4:3, 9:16, and 16:9, assigning each image or video to the bucket with the closest aspect ratio. During training, each data parallel process group samples from one bucket, allowing different buckets across different parallel process groups. We implement longest-side resizing to maximally preserve the original content information described in the prompt. For batch processing, we apply reflection padding to missing pixels and supply the padding mask to the diffusion backbone, enabling precise control during inference.
混合精度训练。我们维护模型权重的两份副本:一份为 BF16,另一份为 FP32。在前向和反向传播过程中,使用 BF16 权重以提高训练效率,梯度和激活值也以 BF16 格式存储。在参数更新时,权重在 FP32 下更新以确保数值稳定性。更新后的 FP32 参数随后被复制并转换为 BF16 用于下一次迭代。为进一步稳定训练,我们将公式 5 中的去噪分数匹配损失缩放 10 倍。我们还发现,AdamW 中较低的 betas 和 eps 系数能显著减少损失尖峰。在我们的 14B 扩散模型训练中,很少遇到损失尖峰,且没有不可恢复的损失尖峰。
Mixed-precision training. We maintain two copies of the model weights: one in BF16 and another in FP32. During the forward and backward passes, the BF16 weights are used to improve training efficiency, resulting in gradients and activations also in BF16 format. For parameter updates, the weights are updated in FP32 to ensure numerical stability. The updated FP32 parameters are then copied and cast to BF16 for the next iteration. To further stabilize training, we scale the loss of denoising score matching in Eq. 5 by a factor of 10. We also find that lower betas and eps coefficients in AdamW significantly reduce loss spikes. For our 14B diffusion model training, we rarely encountered loss spikes, and there were no non-recoverable loss spikes.
文本条件。对于我们的 Text2World 模型,我们采用 T5-XXL 作为文本编码器。我们将 T5 嵌入零填充至固定序列长度 512。为了增强文本上下文对齐,我们采用无分类器引导[引用]。与先前随机丢弃文本嵌入的工作不同,我们省略了此步骤,因为推理时负提示的有效性。值得注意的是,作为文本到图像生成器,我们的模型即使在没有引导的情况下也能生成高保真图像,我们将此能力归因于高质量的训练数据集。虽然无分类器引导通常促进模式寻求行为以获取更优的视觉内容,但我们发现仔细的数据选择也能达到类似效果。然而,对于视频生成,缺乏可比的高质量数据导致在低引导设置下结果欠佳。因此,在视频生成任务中需要更高的引导值才能产生令人满意的内容。
Text conditioning. For our Text2World models, we employ T5-XXL as the text encoder. We zero-pad T5 embeddings to maintain a fixed sequence length of 512. To enhance text-context alignment, we adopt classifier-free guidance [reference]. Unlike prior works that randomly zero out text embeddings, we omit this step due to the effectiveness of negative prompts during inference. Notably, as a text-to-image generator, our model excels in generating high-fidelity images even without guidance, a capability we attribute to the high-quality training dataset. While classifier-free guidance typically promotes mode-seeking behavior for preferred visual content, we find that careful data selection achieves a similar effect. However, for video generation, the lack of comparable high-quality data leads to suboptimal results under low guidance settings. Consequently, higher guidance values are required to produce satisfactory content in video-generation tasks.
图像和视频条件。我们扩展了 Text2World 模型,构建支持图像和视频条件的 Video2World 模型,通过将先前帧纳入生成过程。具体来说,条件帧与生成帧沿时间维度拼接。为了提高推理时对输入帧变化的鲁棒性,我们在训练中对条件帧引入增强噪声。该增强噪声的 sigma 值从\(P_{\text{mean}}=-3.0,P_{\text{std}}=2.0\)采样。此外,扩散模型的输入沿通道维度与一个二进制掩码拼接,该掩码用于区分条件帧和生成帧。损失函数排除条件帧位置的贡献,仅关注生成输出。为了提高泛化能力,我们在训练中随机变化条件帧的数量。在推理时,模型可以灵活地以单个条件帧(图像)或多个先前帧作为输入。
Image and video conditioning. We extend our Text2World models to build Video2World models that support image and video conditioning by incorporating previous frame(s) into the generation process. Specifically, the conditional frame(s) are concatenated with the generated frames along the temporal dimension. To improve robustness against variations in input frame(s) during inference, we introduce augmented noise to the conditional frames during training. The sigma value for this augmented noise is sampled with \(P_{\text{mean}}=-3.0,P_{\text{std}}=2.0\) . Additionally, the input to the diffusion model is concatenated along the channel dimension with a binary mask that distinguishes conditional frames from generated frames. The loss function excludes contributions from the locations of conditional frames, focusing exclusively on the generated output. To improve generalization, we randomly vary the number of conditional frames during training. During inference, the model can flexibly operate with either a single conditional frame (image) or multiple previous frames as input.
在此,我们概述了使我们的扩散世界模型(WFM)能够高效扩展的技术。我们分析了模型的内存需求,讨论了并行化策略,并将我们的训练设置与其他视频扩散模型和最先进的大语言模型(LLM)进行了比较。
Here, we outline the techniques that enable efficient scaling of our diffusion WFMs. We analyze the memory requirements of our models, discuss parallelism strategies, and compare our training setup against other video diffusion models and state-of-the-art LLMs.
共享输入被存储。
The shared input is stored.
查询 \(Q\) 和键 \(K\) 被存储。
The query \(Q\) and key \(K\) are stored.
归一化后的查询 \(Q\) 和键 \(K\) 被重新计算。
The normalized query \(Q\) and key \(K\) are recomputed.
注意力分数( \(A=Q@K^{T}\) )被重新计算。
The attention scores ( \(A=Q@K^{T}\) ) are recomputed.
存储值 \(V\)。重新计算归一化的注意力权重(\(A^{\prime}=\text{Softmax}(A)\))。
The value \(V\) is stored. The normalized attention weights ( \(A^{\prime}=\text{Softmax}(A)\) ) are recomputed.
在交叉注意力中,仅计算查询 \(Q\);键 \(K\) 的序列长度短得多,因此可以忽略不计。
In cross-attention, only query \(Q\) is counted; key \(K\) has much shorter sequence length and is thus negligible.
在交叉注意力中,值 \(V\) 的序列长度短得多,因此可以忽略不计。
In cross-attention, the value \(V\) has much shorter sequence length and is thus negligible.
从 GELU 重新计算输入。
The input is recomputed from GELU.
从 LayerNorm 重新计算输入。
The input is recomputed from LayerNorm.
内存需求。消耗 GPU 内存的四个主要组件是:
Memory requirements. The four major components that consume the GPU memory are:
模型参数:每个参数 10 字节。我们的混合精度训练将模型参数同时以 FP32 和 BF16 存储,并附带 FP32 格式的指数移动平均(EMA)权重。
Model parameters: 10 bytes per parameter. Our mixed precision training stores model parameters in both FP32 and BF16, alongside Exponential Moving Average (EMA) weights in FP32.
梯度:每个参数 2 字节。我们以 BF16 存储梯度。
Gradients: 2 bytes per parameter. We store the gradients in BF16.
优化器状态:每个参数 8 字节。我们使用 AdamW 作为优化器,并以 FP32 存储优化器状态(即一阶和二阶矩)。
Optimizer states: 8 bytes per parameter. We use AdamW as our optimizer and store the optimizer states (i.e., first and second moments) in FP32.
激活值:\((2\times\text{层数}\times 15\times\text{序列长度}\times\text{批大小}\times\text{模型维度})\) 字节。我们以 BF16 存储激活值。表 13 提供了网络内主要操作所存储激活值的详细信息。为了优化内存使用,我们实现了选择性激活检查点,对内存受限的层(如归一化函数)重新计算激活值。
Activations: \((2\times\text{number\_of\_layers}\times 15\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model})\) bytes. We store the activations in BF16. Table 13 provides details of the stored activations for major operations within the network. To optimize memory usage, we implement selective activation checkpointing, recomputing activations for memory-limited layers such as normalization functions.
例如,我们的 14B 模型(Cosmos-Predict1-14B-Text2World)在模型参数、梯度和优化器状态方面需要约 280 GB 内存,而在高分辨率预训练期间,激活内存还需 310 GB。鉴于 NVIDIA H100 GPU 的 80GB HBM3 限制,我们采用全分片数据并行(FSDP)和上下文并行(CP)来将内存需求分布到多个 GPU 上。
For instance, our 14B model (Cosmos-Predict1-14B-Text2World) requires approximately 280 GB for model parameters, gradients, and optimizer states, alongside 310 GB for activations during high-resolution pre-training. Given the 80GB HBM3 limit of NVIDIA H100 GPUs, we employ Fully Sharded Data Parallelism (FSDP) and Context Parallelism (CP) to distribute memory demands across multiple GPUs.
全分片数据并行(FSDP)。FSDP 通过将模型参数、梯度和优化器状态分片到各设备上来提高内存效率。它仅在计算需要时收集参数,并在之后释放。与标准数据并行(在设备间复制参数)不同,FSDP 分布参数、梯度和优化器状态,每个设备仅管理其分片。这种方法将内存使用量降至最大的临时未分片参数集及其参数、梯度和优化器状态分片。在我们的实现中,我们为 7B 模型使用分片因子 32,为 14B 模型使用 64,以平衡内存和通信延迟。
Fully Sharded Data Parallelism (FSDP). FSDP improves memory efficiency by sharding model parameters, gradients, and optimizer states across devices. It gathers parameters only when needed during computation and releases them afterward. Unlike standard data parallelism, which duplicates parameters across devices, FSDP distributes parameters, gradients, and optimizer states, with each device managing only its shard. This approach minimizes memory usage to the largest temporarily unsharded parameter set alongside its shard of parameters, gradients, and optimizer states. For our implementation, we utilize a sharding factor of 32 for the 7B model and 64 for the 14B model to balance memory and communication latency.
上下文并行(CP)。将 Transformer 扩展到长上下文场景会带来 FLOPs 和激活内存增加的挑战。CP 通过将计算和激活分布到多个 GPU 上来解决这些挑战。其工作原理是将查询 \(Q\) 和键值 \((K,V)\) 沿序列维度分割成 CP_SIZE 块,其中 CP_SIZE 是 CP 组内的 GPU 数量。每个 GPU 处理 \(Q\) 的一个块,并使用同一 CP 组中存储的 \((K,V)\) 块迭代累积部分注意力输出。CP 的不同实现使用不同的通信原语,包括全收集、点对点(P2P)和全对全。我们采用 TransformerEngine 中的 P2P 变体,通过在 GPU 之间传输 \((K,V)\) 块的同时处理注意力,实现计算与通信的重叠。当块大小选择恰当时,这种重叠能有效隐藏数据传输延迟。我们将 CP 组组织在 NVLink 连接的 GPU 内,并将 CP 秩与 FSDP 秩重叠以实现最佳利用率。对于上下文较短的图像迭代,禁用 CP 以提高吞吐量。交叉注意力层不使用 CP,因为 \((K,V)\) 的序列长度较短,导致计算量不足以掩盖通信延迟。
Context Parallelism (CP). Scaling transformers for long-context settings introduces challenges with increased FLOPs and activation memory. CP addresses these challenges by distributing computation and activations across multiple GPUs. It works by splitting both the query \(Q\) and the key-value \((K,V)\) along their sequence dimensions into CP_SIZE chunks, where CP_SIZE is the number of GPUs within a CP group. Each GPU processes one chunk of \(Q\) and iteratively accumulates partial attention outputs using blocks of \((K,V)\) stored in the same CP group. Different implementations of CP utilize different communication primitives, including all-gather , P2P , and all-to-all . We employ the P2P variant from TransformerEngine , which overlaps computation and communication by transferring \((K,V)\) blocks between GPUs while simultaneously processing attention. When block sizes are carefully chosen, this overlap effectively hides data transfer latency. We organize CP groups within NVLink-connected GPUs and overlap CP ranks with FSDP ranks for optimal utilization. For image iterations with shorter contexts, CP is disabled to improve throughput. Cross-attention layers do not use CP due to the shorter sequence lengths of \((K,V)\) , which results in insufficient computation to mask communication latency.
以 Cosmos-Predict1-14B 为例,采用分片因子为 64 的 FSDP 可将参数、梯度和优化器状态的内存需求从 280 GB 降至约 \(280\mathbin{/}64\approx 4\) GB/GPU。类似地,采用 \(\text{CP\_SIZE}=8\) 的 CP 可将激活内存从 310 GB 降至约 \(310\mathbin{/}8\approx 40\) GB/GPU。需要注意的是,这些计算是低估的;实际上,分词器和未分片参数会消耗额外内存。CP 中通信与计算的重叠也要求每个 GPU 保留多个 \((K,V)\) 块。
Using Cosmos-Predict1-14B as an instance, employing FSDP with a sharding factor of 64 reduces memory requirements for parameters, gradients, and optimizer states, bringing them down from 280 GB to approximately \(280\mathbin{/}64\approx 4\penalty\ \text{GB per GPU}\) . Similarly, employing CP with \(\text{CP\_SIZE}=8\) decreases activation memory from 310 GB to roughly \(310\mathbin{/}8\approx 40\penalty\ \text{GB per GPU}\) . It is important to note that these calculations are underestimations; in practice, additional memory is consumed by the tokenizer and unsharded parameters. Overlapping communication and computation in CP also necessitates each GPU to retain multiple chunks of \((K,V)\) .
与其他视频生成模型的比较。与 HunyuanVideo 和 MovieGen 中提出的方法相比,我们的并行策略有意简化,那些方法采用了张量并行(TP)及其扩展序列并行(SP)。尽管排除了 TP/SP,我们的设置仍实现了相当的模型 FLOPs 利用率(MFU)。虽然 TP/SP 在某些场景下仍有价值,例如更大的模型或替代网络拓扑,但详细的权衡分析留待未来工作。
Comparison with other video generative models. Our parallelism strategy is deliberately streamlined compared to approaches outlined in HunyuanVideo and MovieGen , which incorporate Tensor Parallelism (TP) and its extension, Sequence Parallelism (SP). Despite excluding TP/SP, our setup achieves comparable Model FLOPs Utilization (MFU). While TP/SP remains valuable in certain scenarios, such as larger models or alternative network topologies, a detailed analysis of tradeoffs is left for future work.
与大型语言模型的比较。与通常以较短上下文长度进行预训练的 LLM 不同,长上下文设置会因自注意力机制的二次方成本而显著增加 FLOPs。虽然 LLM 的 FLOPs 通常按\(6\times\text{seq\_len}\times P\)计算,其中\(P\)为参数数量,但我们注意到该公式对我们的扩散世界模型并不准确。我们在表 13 中提供了每个关键操作的前向传播 FLOPs。
Comparison with large language models. Unlike LLMs, which are typically pre-trained with shorter context lengths, long-context settings significantly increase FLOPs due to the quadratic cost of self-attention. While FLOPs for LLMs are commonly calculated as \(6\times\text{seq\_len}\times P\) , where \(P\) is the number of parameters , we note that this formula is inaccurate for our diffusion WFMs. We provide the forward pass FLOPs of each key operation in table 13.
在训练过程中,我们的世界模型(WFMs)使用详细的视频描述作为输入文本提示,以生成高质量的视频。然而,在推理过程中,用户提示的长度、结构和风格可能各不相同,且往往简短得多。为了弥合训练与推理文本提示之间的差距,我们开发了一个提示词上采样器,将原始输入提示转换为更详细、更丰富的版本。它可以通过添加更多细节并保持一致的描述结构来改进提示,从而产生更高质量的输出。
During training, our WFMs use detailed video descriptions as input text prompts to produce high-quality videos. However, during inference, user prompts may vary in length, structure, and style, often being much shorter. To bridge this gap between training and inference text prompts, we develop a prompt upsampler to transform original input prompts into more detailed and enriched versions. It can improve the prompts by adding more details and maintaining a consistent description structure, which leads to higher quality output.
提示词上采样器的主要要求如下:
The main requirements for the prompt upsampler are:
对输入提示的保真度:上采样后的提示必须忠实保留原始用户输入的关键元素,包括主要角色、动作或运动、关键属性以及整体意图。
Fidelity to the input prompts: The upsampled prompt must faithfully preserve the key elements of the original user input, including the main characters, actions or motions, key attributes, and overall intent.
与训练分布对齐:上采样后的提示应在长度、语言结构和风格上接近世界模型(WFMs)训练提示的分布。
Alignment with training distribution: The upsampled prompt should closely resemble the distribution of training prompts of WFMs in terms of length, language structure, and style.
增强视觉细节:上采样后的提示应设计为引导世界模型(WFMs)生成更准确的图像。
Enhanced visual details: The upsampled prompt should be designed to prompt the WFMs to generate more accurate imagery.
用于 Text2World 模型的提示上采样器。我们微调 Mistral-NeMo-12B-Instruct 来构建我们的提示上采样器。为了获得配对数据,即模拟用户输入的短提示和反映训练提示分布的对应长提示,我们使用 VLM 基于训练长提示和对应视频生成短描述。这种长到短的数据创建策略在以下两方面有效:(1) 保留 WFM 详细训练提示中的真实视频内容和分布;(2) 确保短提示和长提示之间的一致性。由此产生的提示上采样器被称为 Cosmos-UpsamplePrompt1-12B-Text2World。
Prompt upsampler for Text2World model. We fine-tune Mistral-NeMo-12B-Instruct to build our prompt upsampler. To obtain paired data, that is, short prompts simulating user input and the corresponding long prompts reflecting the distribution of training prompts, we use a VLM to generate short captions based on our training long prompts and corresponding videos. This long-to-short data creation strategy is effective in (1) preserving the authentic video content and distribution from detailed training prompts of WFMs and (2) ensuring fidelity between the short and long prompts. The resulting prompt upsampler is termed Cosmos-UpsamplePrompt1-12B-Text2World.
用于 Video2World 模型的提示上采样器。对于 Video2World 模型,输入由视频条件和用户文本提示组成。为了增强用户提示,我们利用开源 VLM Pixtral-12B,结合零样本提示工程,将提示上采样为同时考虑视频条件和用户提示的详细描述。我们发现原版 Pixtral-12B 模型开箱即用效果良好,因此没有继续进行上述类似的微调。
Prompt upsampler for Video2World model. For the Video2World model, the input consists of video conditions and a user text prompt. To enhance the user prompt, we utilize an open-source VLM, Pixtral-12B, combined with zero-shot prompt engineering, to upsample the prompt into a detailed description that considers both the video conditions and the user prompt. We found the vanilla Pixtral-12B model works well out of the box and did not proceed to perform a similar fine-tuning described above.
在图 12 中,我们展示了由 Cosmos-Predict1-7B-Text2World 和 Cosmos-Predict1-14B-Text2World 模型生成的定性结果。两个模型都能生成具有高视觉质量、良好运动动态和文本对齐的视频。与 7B 模型相比,14B 模型能够生成捕捉更复杂视觉细节和精细运动的视频。
In Fig. 12, we present qualitative results generated by our Cosmos-Predict1-7B-Text2World and Cosmos-Predict1-14B-Text2World models. Both models produce videos of high visual quality, motion dynamics, and text alignment. Compared to the 7B model, the 14B model is able to generate videos capturing more complex visual details and intricate motions.
我们在图 13 中展示了 Video2World 7B 和 14B 模型生成的视频。Video2World 模型支持图像和视频条件输入,并能以自回归方式生成扩展视频。如图 13 所示,我们的 Video2World 模型生成的视频具有照片级真实感,运动动态和视觉保真度良好。同样,14B 模型在场景丰富度和运动稳定性方面生成的视频更优。
We show generated videos from Video2World 7B and 14B models in Fig. 13. The Video2World models support both image and video conditioning and can generate extended videos in an autoregressive manner. As demonstrated in Fig. 13, our Video2World models produce photorealistic videos with good motion dynamics and visual fidelity. The 14B model, again, generates better videos in terms of scene richness and motion stability.
在自回归世界基础模型中,我们将世界模拟生成表述为类似于语言建模的下一个词预测任务。首先,我们使用第 4 节中介绍的 Cosmos 离散分词器将视频转换为离散视频令牌序列 \(\mathcal{V}=\{v_{1},v_{2},\dots,v_{n}\}\)。然后,我们训练一个 Transformer 解码器,以过去的视频令牌为上下文预测下一个视频令牌,类似于大型语言模型(LLM)。具体来说,训练目标是最小化以下负对数似然(NLL)损失:
In autoregressive WFMs, we formulate world simulation generation as a next-token prediction task similar to language modeling. We start by converting a video into a sequence of discrete video tokens \(\mathcal{V}=\{v_{1},v_{2},\dots,v_{n}\}\) using the Cosmos Discrete Tokenizer introduced in Section 4. Then we train a Transformer decoder to predict the next video token using past video tokens as context, similar to large language models (LLMs). Specifically, the training objective is to minimize the following negative log-likelihood (NLL) loss:
其中,预测下一个视频令牌 \(v_{i}\) 的条件概率 \(P\) 由参数为 \(\Theta\) 的 Transformer 解码器建模。
where the conditional probability \(P\) of the predicted next video token \(v_{i}\) is modeled by a Transformer decoder with parameters \(\Theta\).
我们的自回归式 WFM 架构如图 14 所示。我们对标准 Transformer 模型架构进行了若干修改,以适应视频生成任务,包括添加 1) 3D 感知位置嵌入,2) 交叉注意力机制以实现文本输入从而更好地控制,以及 3) QK 归一化。
Our autoregressive-based WFM architecture is illustrated in Fig. 14. We make several modifications to the standard transformer model architecture tailored for our video generation task, including adding 1) 3D-aware positional embeddings, 2) cross-attention to enable textual inputs for better control, and 3) QK-Normalization.
3D 位置嵌入。与我们的扩散式 WFM(第 5.1.2 节)类似,我们引入了两种互补的位置嵌入机制:用于相对位置的 3D 分解旋转位置嵌入(RoPE)和用于绝对坐标的 3D 分解绝对位置嵌入(APE)。这些机制协同工作,在整个网络中提供全面的空间和时间信息。
3D positional embeddings. Similar to our diffusion-based WFM (Section 5.1.2), we incorporate two complementary positional embedding mechanisms: 3D factorized Rotary Position Embedding (RoPE) for relative positions and 3D factorized absolute positional embedding (APE) for absolute coordinates. These mechanisms work in concert to provide comprehensive spatial and temporal information throughout the network.
3D 旋转位置嵌入(RoPE)。我们将 3D RoPE 应用于模型,以编码时间、高度和宽度维度上的相对位置信息。在训练过程中,我们采用多阶段训练策略,视频序列长度随训练进度而增加。为了适应不断变化的时间长度,我们使用 YaRN,这是一种计算高效的技术,旨在扩展 RoPE 的上下文窗口。由于视频序列长度仅沿时间维度增加,我们仅沿时间轴应用 YaRN 扩展。通过使用 YaRN,我们的模型可以外推到比初始训练阶段遇到的上下文长度更长的序列。
3D Rotary Position Embedding (RoPE). We apply 3D RoPE to our model to encode relative positional information across the temporal, height, and width dimensions. During training, we adopt a multi-stage training strategy in which the sequence length of videos increases as the training progresses. To adapt the 3D RoPE to the changing temporal duration, we use YaRN, a compute-efficient technique designed to extend the context window of RoPE. We apply YaRN extension only along the temporal axis as the video sequence length increases only along the temporal dimension. By utilizing YaRN, our model can extrapolate to context lengths longer than those encountered during the initial stages of training.
3D 绝对位置嵌入(APE)。除了 3D RoPE 之外,我们在每个 Transformer 块内加入 3D APE 以补充相对位置编码。该 APE 使用跨时间、高度和宽度维度分解的正弦嵌入来编码位置信息,确保模型感知绝对位置。嵌入在每个阶段直接添加到输入张量中,丰富了 Transformer 的位置上下文。我们发现,结合绝对和相对位置编码可以增强模型性能,降低训练损失,并减少生成视频中的变形伪影。值得注意的是,虽然我们的扩散式 WFM(第 5.1.2 节)采用可学习嵌入,但我们在自回归式 WFM 中采用基于正弦的 APE 嵌入。
3D Absolute Positional Embedding (APE). In addition to 3D RoPE, we incorporate a 3D APE within each transformer block to complement the relative positional encoding. This APE encodes positional information using sinusoidal embeddings factorized across temporal, height, and width dimensions, ensuring the model is aware of absolute positions. The embedding is added directly to the input tensor at each stage, enriching the positional context for the transformer. We find combining absolute and relative positional encodings enhances model performance, reduces training loss, and minimizes morphing artifacts in generated videos. Notably, while our diffusion-based WFM (Section 5.1.2) employs learnable embeddings, we adopt sinusoidal-based embeddings for APE in our autoregressive-based WFM.
词汇表。分词是将输入文本转换为离散 token 序列的关键步骤,这在大型语言模型(LLM)中尤为重要。在 LLM 中,可能的 token 词汇表由 LLM 的分词器(例如,[参考文献] 引入的 tiktoken)决定,该分词器使用字节对编码(BPE)等算法在大型文本语料库上训练得到。
Vocabulary. Tokenization is a crucial step that turns input text into a sequence of discrete tokens in large language models (LLMs). In LLMs, the vocabulary of possible tokens is determined by the LLM's tokenizer (e.g., tiktoken introduced by [reference]) trained on a large corpus of text with algorithms such as Byte Pair Encoding (BPE).
对于我们的自回归模型,我们使用 Cosmos-Tokenize1-DV8 \times 16 \times 16-720p 作为分词器。如第 4 节所述,我们利用有限标量量化(FSQ)将 6 维潜在空间量化为 (8,8,8,5,5,5) 个级别。这种量化导致词汇表大小为 8 \times 8 \times 8 \times 5 \times 5 \times 5 = 64,000。
For our autoregressive models, we use our Cosmos-Tokenize1-DV8 \times 16 \times 16-720p as the tokenizer. As introduced in Section 4, we leverage the Finite-Scalar-Quantization (FSQ) to quantize the 6-dimensional latent space into (8,8,8,5,5,5) levels. This quantization leads to a vocabulary size of 8 \times 8 \times 8 \times 5 \times 5 \times 5 = 64{,}000.
用于文本条件的交叉注意力。除了 Transformer 架构中的自注意力块之外,我们还添加了交叉注意力层,以使模型能够根据输入文本进行条件生成。与基于扩散的 WFM(第 5.1.2 节)类似,交叉注意力应用于 Transformer 模型的特征与从预训练文本编码器(T5-XXL)获得的文本嵌入之间。在我们的实验中,我们在每个自注意力层之后添加交叉注意力块。
Cross-attention for text conditioning. In addition to the self-attention blocks present in the transformer architecture, we add cross-attention layers to enable the model to condition on input text. Similar to diffusion-based WFM (Section 5.1.2), cross-attention is applied between the features of the transformer model and text embeddings obtained from a pre-trained text encoder (T5-XXL). In our experiments, we add cross-attention blocks after every self-attention layer.
查询-键归一化。为了增强训练稳定性,我们引入了查询-键归一化(QKNorm)。QKNorm 通过在计算点积之前对查询(Q)和键(K)向量进行归一化来解决注意力机制中的不稳定性,从而防止 softmax 函数饱和,确保更有效的学习。归一化后,点积由一个可学习参数 \gamma 缩放,而不是固定的 1/\sqrt{d_k}。这种可学习的缩放因子使模型能够自适应地控制注意力分数的幅度,增强灵活性和表达能力。
Query-key normalization. In order to enhance training stability, we incorporate Query-Key Normalization (QKNorm). QKNorm addresses instability in attention mechanisms by normalizing the query (Q) and key (K) vectors before computing their dot product, thereby preventing the softmax function from saturating and ensuring more effective learning. After normalization, the dot product is scaled by a learnable parameter \gamma instead of the fixed 1/\sqrt{d_k}. This learnable scaling factor allows the model to adaptively control the magnitude of the attention scores, enhancing flexibility and expressivity.
Z-loss。为了进一步提高训练稳定性,我们在训练目标中引入了一个称为 z-loss 的稳定项。z-loss 惩罚 logits 偏离零的情况,有效阻止模型生成过大的 logit 值,这些值可能导致数值不稳定或梯度爆炸。z-loss 定义为 logits 的平方和,即 \mathcal{L}_{\text{z-loss}} = \lambda \cdot \sum_i z_i^2。我们发现 z-loss 对于将梯度范数维持在健康范围内至关重要,尤其是在将训练扩展到大量 GPU 节点时。根据经验,我们发现 z-loss 系数 \lambda = 3 \times 10^{-4} 达到了最佳平衡,既能有效稳定训练,又不会对模型性能产生不利影响。
Z-loss. To further improve training stability, we introduce a stabilization term known as the z-loss into our training objective. The z-loss penalizes deviations of the logits from zero, effectively discouraging the model from generating excessively large logit values that could result in numerical instability or gradient explosions. The z-loss is defined as the sum of the squared logits as \mathcal{L}_{\text{z-loss}} = \lambda \cdot \sum_i z_i^2. We found z-loss to be critical in maintaining gradient norms to a healthy range, especially when scaling the training to a large number of GPU nodes. Empirically, we found that the z-loss coefficient \lambda = 3 \times 10^{-4} strikes an optimal balance, effectively stabilizing training without adversely affecting model performance.
本节介绍使我们自回归世界模型(WFM)能够高效扩展(scaling)的技术。我们简要分析模型的内存消耗,讨论并行策略,并将我们的训练设置与其他自回归模型进行比较。
This section describes the techniques that enable efficient scaling of our autoregressive WFMs. We briefly analyze the memory consumption of our models, discuss parallelism strategies, and compare our training setup with other autoregressive models.
内存需求。在训练过程中,GPU 内存主要消耗在:
Memory requirements. During training, GPU memory is mainly consumed by:
模型参数:每个参数 6 字节。我们以 BF16 和 FP32 两种格式存储模型参数。
Model parameters: 6 bytes per parameter. We store the model parameters in both BF16 and FP32.
梯度:每个参数 2 字节。我们以 BF16 格式存储梯度。
Gradients: 2 bytes per parameter. We store the gradients in BF16.
优化器状态:每个参数 8 字节。我们以 FP32 格式存储 AdamW 的一阶和二阶矩。
Optimizer states: 8 bytes per parameter. We store the first and second moments of AdamW both in FP32.
激活值:大约为 \((2\times\text{层数}\times 17\times\text{序列长度}\times\text{批大小}\times\text{模型维度})\) 字节。我们建议读者参阅文献[引用]以了解最先进自回归模型激活值内存的详细分析。
Activations: Approximately \((2\times\text{number\_of\_layers}\times 17\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model})\) bytes. We refer readers to [reference] for a detailed analysis of activation memory of state-of-the-art autoregressive models.
例如,我们的 12B 模型(Cosmos-Predict1-12B)的参数、梯度和优化器状态合计需要约 192 GB 内存。由于这超出了单个 NVIDIA H100 GPU 的 80GB HBM3 容量,我们利用张量并行(TP)及其扩展序列并行(SP)将内存需求和计算分布到多个 GPU 上。
For instance, our 12B model (Cosmos-Predict1-12B) demands approximately 192 GB of memory for its parameters, gradients, and optimizer states combined. As this is beyond a single NVIDIA H100 GPU’s 80GB HBM3 capacity, we leverage tensor parallelism (TP) and its extension, sequence parallelism (SP), to distribute the memory requirements and computation across multiple GPUs.
张量并行(TP)。张量并行(TP)沿输入或输出特征维度分割线性层的权重,选择依据是最小化 GPU 间通信。例如,在一个两层前馈网络中,第一层的权重沿输出特征维度分割,而第二层的权重沿输入特征维度分割。这种安排使得中间激活值可以在本地处理,无需 GPU 间通信。最终输出通过全归约通信合并。采用 TP 后,每个 GPU 仅存储线性层权重的一部分,具体为 \(1/\text{TP\_SIZE}\)。然而,TP 的默认实现仍会在序列维度上复制激活值,用于 LayerNorm 等操作,导致冗余。
Tensor Parallelism (TP). Tensor Parallelism (TP) splits the weights of linear layers along either the input or output feature dimensions, with the choice guided by the goal of minimizing inter-GPU communication. For example, in a two-layer feedforward network, the weights of the first layer are partitioned along the output feature dimension, while those of the second layer are partitioned along the input feature dimension. This arrangement allows intermediate activations to be processed locally without requiring communication between GPUs. The final outputs are then combined using all-reduce communication. By employing TP, each GPU stores only a fraction, specifically \(1/\text{TP\_SIZE}\), of the weights for linear layers. However, the default implementation of TP still replicates activations along the sequence dimension for operations like LayerNorm, resulting in redundancy.
序列并行(SP)。SP 通过沿序列维度进一步划分上下文来扩展张量并行。该方法适用于自注意力层中的 LayerNorm 和 Dropout 等算子,这些算子中序列的每个元素可以独立处理。启用 SP 后,每个 GPU 仅存储激活值的一部分,具体为 \(1/\text{TP\_SIZE}\)。
Sequence Parallelism (SP). SP extends Tensor Parallelism by further partitioning the context along the sequence dimension. This approach is applicable to operators, such as LayerNorm and Dropout in self-attention layers, where each element in the sequence can be processed independently. With SP enabled, each GPU stores only a fraction, specifically \(1/\text{TP\_SIZE}\), of the activations.
与其他自回归模型的比较。与流行的 LLM 相比,我们的模型没有利用诸如 MQA 或 GQA 等节省内存的注意力变体。除此之外,我们的自回归模型特意设计为与 LLM 的架构紧密相似,因为这种对齐提供了灵活性和可扩展性。利用更多并行方式(如上下文并行和流水线并行)进一步扩大模型规模和上下文长度的实验留待未来工作。
Comparison with other autoregressive models. Compared to popular LLMs, our model doesn’t leverage memory-saving attention variants such as MQA or GQA. Otherwise, our autoregressive model is deliberately designed to closely resemble the architecture of LLMs, as this alignment offers flexibility and scalability. Experiments that leverage more parallelisms, such as context parallelism and pipeline parallelism, to further scale up the model sizes and context lengths are left for future works.
我们对自回归 WFM 进行多阶段的预训练。
We perform pre-training of our autoregressive WFMs in multiple stages.
阶段 1:在第一阶段,模型使用视频预测目标进行训练。给定第一帧作为输入条件,模型被训练来预测未来的视频帧。此任务使用 17 帧的上下文长度,即模型以第一帧为输入预测 16 个未来帧。
Stage 1: In the first stage, the model is trained using the video prediction objective. Given the first frame as the input condition, the model is trained to predict future video frames. A context length of 17 frames is used for this task, i.e., the model predicts 16 future frames with the first frame as input.
阶段 1.1:此阶段执行视频预测,但上下文长度增加到 34 帧。我们在时间维度上使用 YaRN 扩展来增加 RoPE 的上下文长度。
Stage 1.1: This stage performs video prediction but with an increased context length of 34 frames. We use the YaRN extension on the temporal dimension to increase the context length of RoPE.
阶段 2:在训练的第二阶段,我们向模型引入文本条件。通过新初始化的交叉注意力层加入文本嵌入。模型以 34 帧的上下文进行训练。为了提高文本到视频的生成能力,模型使用联合图像和视频数据进行训练,如第 5.1.3 节所述。当使用图像批次时,我们使用更大的批次大小,因为图像的上下文长度远小于视频。
Stage 2: In stage 2 of our training, we introduce text conditioning to our model. Text embeddings are incorporated using newly initialized cross-attention layers. The model is trained with a 34-frame context. To improve text-to-video generation ability, the model is trained using joint image and video data as described in section 5.1.3. When image batches are used, we use a larger batch size as the context length for images is much smaller than that of videos.
我们所有的模型都以固定的空间分辨率 \(640\times 1024\) 进行训练。
All our models are trained with a fixed spatial resolution of \(640\times 1024\) .
冷却阶段。预训练之后,我们使用高质量数据执行一个“冷却”阶段,类似于大语言模型的训练实践。在此阶段,我们在高质量图像-视频对上训练时,将学习率线性衰减至 0。冷却阶段持续 30,000 次迭代。
Cooling down. After pre-training, we conduct a “cooling-down” phase with high-quality data, similar to LLM training practices. During this phase, we linearly decay the learning rate to 0 while training on high-quality image-video pairs. The cooling-down phase is carried out over 30,000 iterations.
我们训练了两组基于自回归的世界基础模型(WFM)。首先构建两个基础模型:一个具有 4B 容量,另一个具有 12B 容量。这些是纯粹的下一个视频词元预测器,不接受文本提示作为输入。然后,我们从每个基础模型派生出一个 Video2World 版本,在其中添加交叉注意力层,以利用文本提示输入进行下一个视频词元预测。
We train two sets of autoregressive-based WFMs. We start by building two base models: one with a 4B capacity and the other with a 12B capacity. These are pure next-video token predictors that do not take text prompts as input. We then derive a Video2World version from each of the base models, where we add cross-attention layers to them to leverage text prompt inputs for next video token prediction.
Cosmos-Predict1-4B:一个用于下一个视频词元预测的 4B Transformer 模型。该模型使用多阶段训练目标中的阶段 1 和阶段 1.1 进行训练。
Cosmos-Predict1-4B: a 4B transformer model for next video token prediction. This model is trained using stage 1 and stage 1.1 of the multi-stage training objective.
Cosmos-Predict1-5B-Video2World:一个从我们的 Cosmos-Predict1-4B 派生的 5B Transformer 模型,并额外使用多阶段训练目标中的阶段 2 进行训练。
Cosmos-Predict1-5B-Video2World: a 5B transformer model derived from our Cosmos-Predict1-4B and trained additionally with stage 2 of the multi-stage training objective.
Cosmos-Predict1-12B:一个用于下一个视频词元预测的 12B Transformer 模型。该模型使用多阶段训练目标中的阶段 1 和阶段 1.1 进行训练。
Cosmos-Predict1-12B: a 12B transformer model for next video token prediction. This model is trained using stage 1 and stage 1.1 of the multi-stage training objective.
Cosmos-Predict1-13B-Video2World:一个 13B 的 Transformer 模型,源自 Cosmos-Predict1-12B,并额外使用多阶段训练目标的第二阶段进行训练。
Cosmos-Predict1-13B-Video2World: a 13B transformer model derived from Cosmos-Predict1-12B and trained additionally with stage 2 of the multi-stage training objective.
我们的 Cosmos 自回归世界模型与 LLM 在架构上具有相似性,这使我们能够利用成熟的 LLM 推理优化技术来解决顺序解码的瓶颈。我们实现了键值缓存、张量并行和 torch.compile 的组合,遵循 PyTorch 中的 gpt-fast 实现。
Our Cosmos Autoregressive WFMs share architectural similarities with LLMs, enabling us to leverage established LLM inference optimization techniques to address the sequential decoding bottleneck. We implement a combination of key-value caching, tensor parallelism, and torch.compile, following the gpt-fast implementation in PyTorch.
投机解码。为了进一步加速我们的自回归世界模型,我们应用了 Medusa 投机解码框架。与常见的需要单独草稿模型的投机解码方法或加速有限的免训练方法不同,Medusa 通过额外的解码头扩展 Transformer 主干,以并行预测多个后续词元。然后,它通过拒绝采样验证这些推测的词元。通过缓解一次一个词元处理的瓶颈,推理得以加速。我们展示了 Medusa 技术在视觉自回归加速中的潜力,且不损害生成输出的质量。
Speculative decoding. To further accelerate our autoregressive WFMs, we apply the Medusa speculative decoding framework. Unlike common speculative decoding approaches that require a separate draft model or training-free methods with limited speedup, Medusa extends the transformer backbone with extra decoding heads to predict multiple subsequent tokens in parallel. It then verifies these speculated tokens with rejection sampling. The inference is thus accelerated by alleviating the bottleneck of one-token-at-a-time processing. We demonstrate the potential of the Medusa technique in visual autoregressive acceleration without compromising the quality of generated outputs.
在我们的实现中,我们通过在架构中引入 Medusa 头来微调预训练的自回归世界模型。这些头被策略性地插入到最后一个 Transformer 隐藏状态之后,其中所有主干参数和最终解嵌入层在不同头之间共享。每个 Medusa 头是一个带有 SiLU 激活和残差连接的单层 FFN。我们进一步将多个 Medusa 头的权重矩阵合并为一个统一的 FFN,以最大化词元预测期间的并行性。注意,我们不使用原始 Medusa 论文中的基于树的注意力机制。
In our implementation, we fine-tune our pre-trained autoregressive WFMs by introducing Medusa heads into the architecture. These heads are strategically inserted after the last transformer hidden states, where all backbone parameters and the final unembedding layer are shared across different heads. Each Medusa head is a single-layer FFN with SiLU activation and residual connection. We further merge the weight matrices of multiple Medusa heads into a unified FFN to maximize parallelism during token prediction. Note that we do not use the tree-based attention mechanism from the original Medusa paper.
为了研究我们的自回归世界模型的最佳 Medusa 设置,我们从两个方面进行了深入研究:(1) 微调哪些 Transformer 层;(2) 添加多少个 Medusa 头。对于第一个问题,我们比较了全量微调和选择性层冻结。我们观察到,仅微调 Medusa 头会导致多词元预测效果不佳,而全量微调则会导致质量下降。我们通过实验确定,解冻最后两个 Transformer 层和最终解嵌入层,同时保持主干冻结,可以获得最佳性能。该策略确保我们的 Medusa 训练在实现良好的投机解码准确性的同时,不会遭受灾难性遗忘。
To investigate the optimal Medusa setup for our autoregressive WFMs, we conduct an in-depth study from two aspects: (1) which transformer layers to fine-tune and (2) how many Medusa heads to add. For the first problem, we compare between full fine-tuning and selective layer freezing. We observe that only fine-tuning the Medusa heads gives poor multi-token prediction, while full fine-tuning incurs quality degradation. We empirically identify that unfreezing the last two transformer layers and the final unembedding layer while keeping the backbone frozen yields the best performance. This strategy ensures our Medusa training achieves decent speculative decoding accuracy without suffering from catastrophic forgetting.
为了探索 Medusa 头的最佳数量,我们计算了不同 Medusa 头数量下的模型词元吞吐量和前向传播次数。消融研究在 8 块 H100 GPU 上进行,并在 50 个未见过的 \(640\times 1024\) 分辨率测试视频上评估。表 15 的结果表明,我们的 Medusa 框架可以有效加速推理,对于 4B 模型,词元吞吐量最高提升 \(2.0\times\),前向传播次数减少 \(4.6\times\);对于 5B 模型,词元吞吐量最高提升 \(3.2\times\),前向传播次数减少 \(6.1\times\)。我们表明,虽然更多的 Medusa 头可以减少生成所需的前向传播次数,但可能会降低整体词元吞吐量。我们发现,\(9\) 个 Medusa 头在计算效率和模型性能之间取得了最佳平衡。
To explore the optimal number of Medusa heads, we calculate the model token throughput and forward pass count with different numbers of Medusa heads. The ablation studies are conducted on 8 \(\times\) H100 GPUs and evaluated on 50 unseen test videos of \(640\times 1024\) resolution. The results in table 15 suggest that our Medusa framework can effectively accelerate inference, with up to \(2.0\times\) token throughput and \(4.6\times\) less forward pass for the 4B model, and up to \(3.2\times\) token throughput and \(6.1\times\) less forward pass for the 5B model. We show that though more Medusa heads can reduce the number of forward passes needed to generate, it may slow down the overall token throughput. We find that \(9\) Medusa heads yield the best trade-off between computational efficiency and model performance.
该表报告了不同设置下各种 Cosmos 自回归 WFM 的平均推理时间(秒)和 VRAM 利用率。推理时间以输入单个条件帧生成 32 帧为准。No DD:无扩散解码器的时间。No DD+Medusa:无扩散解码器但有 Medusa 头的时间。With DD:有扩散解码器的时间。With DD+Medusa:有扩散解码器和 Medusa 头的时间。VRAM:视频内存使用量(GB)。
The table reports the average inference time (in seconds) and VRAM utilization of various Cosmos Autoregressive WFMs under different settings. Inference time is reported for generating 32 frames with a single conditioning frame as input. No DD: Time without diffusion decoder. No DD+Medusa: Time without diffusion decoder but with Medusa heads. With DD: Time with diffusion decoder. With DD+Medusa: Time with diffusion decoder and Medusa heads. VRAM: Video RAM usage in gigabytes.
在表 16 中,我们展示了集成 Medusa 的自回归 WFM 的性能分析。该分析在 H100 GPU 上进行,并在 BF16 精度下对 640×1024 分辨率的测试视频进行评估。结果表明,Medusa 实现在不同 GPU 配置下均能持续加速 4B 和 5B 模型的推理。
In Table 16, we show performance analysis of autoregressive WFMs with Medusa integration. This analysis was conducted on H100 GPUs and evaluated on test videos of 640×1024 resolution in the BF16 precision. Results show that the Medusa implementation consistently accelerates inference for both 4B and 5B models under different GPU configurations.
低分辨率适配以实现实时推理。我们通过将模型适配到较低的空间分辨率 320×512 来追求实时推理,这导致每个视频的 token 数量减少。具体来说,我们首先使用目标 Physical AI 领域的视频,在 320p 低分辨率视频上微调离散视频分词器(第 4 节中的 Cosmos-Tokenize1-DV8 × 16 × 16-720p)。然后,我们使用该低分辨率分词器,在目标 Physical AI 领域的 320×512 分辨率视频上微调预训练于 640×1024 分辨率的自回归 WFM(第 5.2.3 节中的 Cosmos-Predict1-4B)。最后,我们将 Medusa 头添加到微调后的低分辨率自回归 WFM 上。
Low-resolution adaptation for real-time inference. We pursue real-time inference by adapting our model to a lower spatial resolution of 320×512, which results in a lower number of tokens per video. Specifically, we first fine-tune the discrete video tokenizer (Cosmos-Tokenize1-DV8 × 16 × 16-720p in Section 4) on 320p low-resolution videos using videos from the target Physical AI domain. Then, we fine-tune our autoregressive WFM that is pre-trained in 640×1024 resolution (Cosmos-Predict1-4B in Section 5.2.3) with this low-resolution tokenizer on videos of 320×512 resolution from the target Physical AI domain. Finally, we add the Medusa heads to the fine-tuned low-resolution autoregressive WFM.
我们使用 torch.compile 的“max-autotune”模式,在 BF16 精度下于 8 × H100 GPU 上进行了推理基准测试,并使用来自目标 Physical AI 领域的 10-FPS 输入视频进行评估。在表 17 中,我们报告了该设置下实现的平均 token 吞吐量和帧生成速度。我们观察到,我们的模型可以在不到 1 秒内生成 10 个视频帧,这表明我们可以实现 10 FPS 的实时视频生成。
We conducted inference benchmarking on 8 × H100 GPUs using torch.compile's “max-autotune” mode in BF16 precision, and evaluated with 10-FPS input videos from the target Physical AI domain. In Table 17, we report the average token throughput and frame generation speed achieved in this setup. We observe that our model can generate 10 video frames in less than 1 second, demonstrating that we can achieve real-time video generation at 10 FPS.
我们的 Cosmos 分词器采用轻量级编码器-解码器架构进行激进压缩,从而减少 WFM 训练所需的词元数量。然而,激进压缩有时会导致视频生成中出现模糊和可见伪影,尤其是在自回归 WFM 设置下,仅用少量整数通过离散词元化来表示丰富的视频。我们采用扩散解码器设计来解决这一局限。具体来说,我们通过微调第 5.1 节中的 Cosmos-Predict1-7B-Text2Video 来构建更强大的分词器解码器。
Our Cosmos tokenizer uses a lightweight encoder-decoder architecture to perform aggressive compression, which reduces the number of tokens for our WFM training. As a result of aggressive compression, it could sometimes lead to blurriness and visible artifacts in video generation, especially in the autoregressive WFM setting, where only a few integers are used to represent a rich video through discrete tokenization. We resort to the diffusion decoder design to address the limitation. Specifically, we build a more powerful tokenizer decoder by fine-tuning Cosmos-Predict1-7B-Text2Video in section 5.1.
图 16 展示了我们如何为自回归 WFM 训练扩散解码器。对于每个训练视频,我们分别使用 Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p 和 Cosmos-Tokenize1-DV8 \(\times\) 16 \(\times\) 16-720p 计算连续词元视频和对应的离散词元视频。我们注意到,Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p 能产生比 Cosmos-Tokenize1-DV8 \(\times\) 16 \(\times\) 16-720p 更高质量的视频输出,这得益于更温和的连续词元化过程和更不激进的压缩方案(\(8\times 8\times 8\) 而非 \(8\times 16\times 16\))。
Fig. 16 illustrates how we train a diffusion decoder for our autoregressive WFMs. For each training video, we use Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p and Cosmos-Tokenize1-DV8 \(\times\) 16 \(\times\) 16-720p to compute a continuous token video and a corresponding discrete token video, respectively. We note that Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p can produce higher quality video outputs than Cosmos-Tokenize1-DV8 \(\times\) 16 \(\times\) 16-720p thanks to the more gentle continuous tokenization process and the less aggressive compression scheme ( \(8\times 8\times 8\) instead of \(8\times 16\times 16\) ).
离散词元视频被视为 Cosmos-Predict1-7B 模型去噪器的条件输入。为了计算条件输入,我们首先通过可学习的词表嵌入层将离散词元视频中的每个离散词元嵌入到 16 维向量中。然后,我们在 \(x\) 和 \(y\) 方向上将嵌入上采样 \(2\times\),使得条件输入与来自连续词元视频的去噪器噪声输入尺寸相同。我们沿通道维度将噪声连续输入与条件输入拼接,作为扩散去噪器的输入。去噪器的第一层扩展了通道维度以适应新的输入形状。我们通过去除添加的噪声来微调更新后的 Cosmos-Predict1-7B。由于离散词元视频未被噪声破坏,去噪器学会利用条件输入中的信息进行去噪。结果是为分词器提供了一个更高质量的解码器,通过求解逆向扩散问题来解码离散词元。
The discrete token video is treated as the conditional input to the denoiser of the Cosmos-Predict1-7B model. To compute the conditional input, we first embed each discrete token of the discrete token video into a 16-dimensional vector based on a learnable vocabulary embedding layer. We then upsample the embedding \(2\times\) along the \(x\) and \(y\) directions so that the conditional input will be of the same size as the noisy input to the denoiser from the continuous token video. We concatenate the noisy continuous inputs with the conditional inputs along the channel dimension, which becomes the input to the diffusion denoiser. The first layer of the denoiser is channel-dimension expanded to accommodate the new input shape. We fine-tune the updated Cosmos-Predict1-7B by removing the added noise. As the discrete token video is not noise-corrupted, the denoiser learns to leverage the residing information in the conditional input for denoising. The result is a higher-quality decoder for the tokenizer that decodes the discrete token by solving a reserve diffusion problem.
图 16 展示了推理过程。来自自回归 WFM 的输出离散词元视频(在 \(8\times 16\times 16\) 离散压缩下)通过两步解码为视频。首先,我们基于自回归 WFM 输出,展开条件去噪器生成连续词元视频(在 \(8\times 8\times 8\) 连续压缩下)。接下来,连续词元视频由 Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p 解码,生成最终的 RGB 视频。
Fig. 16 illustrates the inference. The output discrete token video (under \(8\times 16\times 16\) discrete compression) from our autoregressive WFM is decoded into a video through two steps. First, we roll out the conditional denoiser to generate a continuous token video (under \(8\times 8\times 8\) continuous compression) based on the autoregressive WFM output. Next, the continuous token video is decoded by Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p to produce the resulting RGB video.
在图 17 中,我们展示了不同模型规模的自回归 WFM 的定性结果。在无提示设置下,比较 Cosmos-Predict1-4B 和 Cosmos-Predict1-12B 模型,我们观察到 12B 模型生成的视频具有更好的运动效果和更清晰的细节。类似地,在提示设置下,比较 Cosmos-Predict1-5B-Video2World 和 Cosmos-Predict1-13B-Video2World 表明,13B 模型比 5B 模型获得了更好的运动效果。
In Fig. 17, we show qualitative results of our autoregressive WFMs using different model sizes. In the unprompted setting, comparing Cosmos-Predict1-4B and Cosmos-Predict1-12B models, we observe that the 12B model generates videos with better motion and sharper details. Similarly, in the prompted setting, comparing Cosmos-Predict1-5B-Video2World and Cosmos-Predict1-13B-Video2World reveals that the 13B model achieves better motion than the 5B model.
在图 18 中,我们展示了使用扩散解码器所获得的增强效果。自回归模型的输出模糊,主要是由于离散分词器中的有损压缩。使用扩散解码器可以在保留内容的同时增强细节。
In Fig. 18, we show the enhancements obtained when using the diffusion decoder. The outputs of the autoregressive model are blurry mainly due to the lossy compression in our discrete tokenizer. The use of the diffusion decoder can enhance details while preserving the content.
我们通过实验发现,基于自回归的 Text2World WFM 的输出并未因第 5.1.5 节中讨论的提示上采样器而得到改善。我们推测这可能是因为这些 WFM 在大部分训练过程中仅使用纯视频生成任务进行预训练,它们没有被足够强制地利用文本输入。
We empirically find that the outputs of the autoregressive-based Text2World WFMs do not improve with upsampled prompts from the prompt upsampler discussed in Section 5.1.5. We hypothesize this is possibly due to the fact that these WFMs are pretrained with pure video generation tasks for most of the training. They are not forced hard enough to leverage text inputs.
在我们的自回归世界基础模型(WFM)生成的视频中,一个显著的失败案例是物体意外地从下方出现。图 19 展示了这一问题的一个示例。为了了解我们模型的失败率,我们进行了一项系统性研究,创建了一个包含 100 个物理 AI 输入的自回归 WFM 评估集。我们使用两种输入模式——图像(单帧)条件和视频(9 帧)条件——用我们所有模型生成视频。对于所有生成的视频,我们手动检查失败案例,并在表 18 中报告失败率。我们观察到,较小的模型 Cosmos-Predict1-4B 和 Cosmos-Predict1-5B-Video2World 在单帧条件下显示出较高的损坏率,而较大的模型 Cosmos-Predict1-12B 和 Cosmos-Predict1-13B-Video2World 则更为稳健。对于所有模型,使用 9 帧视频条件的生成是稳定的,失败率低于 2%。
One notable failure case observed in the generated videos of our autoregressive WFMs is objects unexpectedly appearing from below. Fig. 19 illustrates an example of this issue. To understand the failure rate of our models, we conduct a systematic study by creating an evaluation set of 100 Physical AI inputs to our autoregressive WFMs. We generate videos with all our models using two input modes—image (single-frame) conditioning and video (9-frame) conditioning. For all generated videos, we manually inspect the failure cases and report the failure rate in Table 18. We observe that the smaller models Cosmos-Predict1-4B and Cosmos-Predict1-5B-Video2World show a higher corruption rate in single-frame conditioning, while the larger models Cosmos-Predict1-12B and Cosmos-Predict1-13B-Video2World are more robust. Generation with 9-frame video conditioning is stable for all models, with a failure rate lower than 2%.
预训练的世界模型(WFM)是视觉世界模拟的通才。其能力应从多个方面进行衡量。在此,我们从两个方面评估我们的模型。首先,我们评估生成视频的 3D 一致性。理想的世界模型应从几何上合理的 3D 世界生成视频模拟。其次,我们评估生成视频的物理对齐程度。我们计算渲染的动态在多大程度上符合物理定律。世界模型的评估是一项非常不平凡的任务。我们承认还有其他几个重要方面需要评估。我们将更全面的评估留作未来的工作。
Pre-trained WFMs are generalists of visual world simulation. Their capabilities should be measured across multiple aspects. Here, we evaluate our models on two aspects. First, we evaluate the 3D consistency of the generated videos. An ideal WFM should generate video simulations from geometrically plausible 3D worlds. Second, we evaluate the physics alignment of the generated videos. We calculate how well the rendered dynamics adhere to the laws of physics. Evaluation of WFMs is a highly nontrivial task. We acknowledge that there are several other important aspects required for evaluation. We leave a more comprehensive evaluation as future work.
世界基础模型(WFM)旨在通过视频生成来模拟三维世界,因此评估生成视频与视觉世界三维结构的一致性至关重要。除了看起来逼真之外,生成的视频还应随时间保持与场景物理原理的一致性,这是下游物理 AI 应用的关键要求。
WFMs are designed to simulate 3D worlds through video generation, and it is essential to evaluate how well the generated videos are consistent with the 3D structure of the visual world. In addition to appearing realistic, the generated videos should maintain coherence with the physical principles of scenes through time, a key requirement for downstream Physical AI applications.
测试数据与基线模型。我们聚焦于静态场景,以便利用基于多视图几何的现有工具有效衡量视频的三维一致性。我们从 RealEstate10K 数据集的测试集中随机选取了 500 个视频构建数据集。此外,我们使用专有的视觉语言模型(VLM)为视频添加字幕,获得描述静态场景的文本提示,从而在计算指标时无需考虑场景运动。我们以 VideoLDM 作为基线方法进行比较。
Test data and baseline model. We focus on the scenario of static scenes in order to effectively measure 3D consistency of videos with existing tools based on multi-view geometry. We curate a dataset of 500 videos randomly chosen from the test set of the RealEstate10K dataset. We additionally caption the videos using a proprietary VLM to obtain text prompts that describe the videos as static scenes, so one does not need to consider scene motions for metric computation. We compare against VideoLDM as the baseline method.
指标。生成的视频本质上是底层三维视觉世界的二维投影。我们设计了以下指标来衡量生成视频的三维一致性。
Metrics. Generated videos are effectively 2D projections of the underlying 3D visual worlds. We design the following metrics to measure the 3D consistency of generated videos.
几何一致性。我们通过量化对极几何约束的满足程度来评估生成世界的三维一致性,包括 Sampson 误差以及相机位姿估计算法在生成视频上的成功率。
Geometric consistency. We evaluate the 3D consistency of our generated worlds by quantifying how the epipolar geometry constraints are satisfied, including the Sampson error and the success rate of camera pose estimation algorithms on the generated videos.
视图合成一致性。我们评估世界基础模型在插值新视角下合成图像的能力,同时保持与底层三维结构的一致性。
View synthesis consistency. We evaluate the ability of world foundation models to synthesize images at interpolated novel viewpoints while maintaining coherence with the underlying 3D structure.
Sampson 误差是从一个兴趣点到另一视图对应极线的距离的一阶近似。给定一个帧对中的\(N\)个点对应(以齐次坐标表示)\(\{\left(\bar{\mathbf{x}}_{i},\bar{\mathbf{y}}_{i}\right)\}_{i=1}^{N}\),我们定义 Sampson 误差为
The Sampson error is the first-order approximation of the distance from one interest point to its corresponding epipolar line in another view. Given \(N\) point correspondences (represented in homogeneous coordinates) \(\{\left(\bar{\mathbf{x}}_{i},\bar{\mathbf{y}}_{i}\right)\}_{i=1}^{N}\) in a given frame pair, we define the Sampson error as
其中\(\mathbf{F}\)是从对应点估计的基础矩阵。我们使用误差函数的平方根版本,使度量以像素为单位更直观。我们结合 SuperPoint 和 LightGlue 来检测和匹配帧对的关键点对应,并使用 OpenCV 的 8 点 RANSAC 算法估计\(\mathbf{F}\)。我们将平均误差按帧的对角线长度归一化,相对于\(960\times 540\)的画布。
and \(\mathbf{F}\) is the fundamental matrix estimated from the correspondences. We use the square root version of the error function to make the metric more intuitive in pixel units. We use a combination of SuperPoint and LightGlue to detect and match keypoint correspondences from a frame pair and estimate \(\mathbf{F}\) using OpenCV’s 8-point RANSAC algorithm. We normalize the average error by the diagonal length of the frame with respect to a \(960\times 540\) canvas.
我们还通过生成视频自合成新视角的能力来评估其三维一致性。遵循新视角合成文献的常见做法,我们每隔 8 帧留出测试帧,并使用 Nerfstudio 库的默认设置,用其余训练帧拟合 3D 高斯泼溅模型。我们报告峰值信噪比(PSNR)、结构相似性(SSIM)和 LPIPS 作为量化合成测试视图质量的指标。
We also evaluate 3D consistency of a generated video with its ability to self-synthesize novel viewpoints. Following the common practice of novel view synthesis literature, we hold out every 8 frames as the test frames and fit a 3D Gaussian splatting model with the rest of the training frames using the default settings from the Nerfstudio library. We report the Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and LPIPS as the metrics to quantify the quality of the synthesized test views.
结果。我们在表 19 中展示了定量评估结果。Cosmos WFM 在几何一致性和视图合成一致性方面均显著优于我们的基线模型。Cosmos WFM 的兴趣点不仅更具三维一致性,而且相机位姿估计成功率也明显更高,这反映了整体质量的提升和三维一致性的增强,甚至达到了真实世界视频的水平。在成功估计相机位姿的案例中,合成的留出视图在所有图像合成指标上均表现出更高的质量。这些结果凸显了我们的 Cosmos WFM 生成三维一致视频的能力,使其成为有效的世界模拟器。
Results. We present the quantitative evaluation results in table 19. The Cosmos WFMs achieve significantly better 3D consistency than our baseline model in terms of both geometric and view synthesis consistency. Not only are the interest points from Cosmos WFMs more 3D-consistent, but the camera pose estimation success rate is also notably higher, reflecting both improved overall quality and enhanced 3D consistency, even reaching the level of real-world videos. Among the cases where camera poses were successfully estimated, the synthesized held-out views demonstrate higher quality across all image synthesis metrics. These results highlight the capability of our Cosmos WFMs to generate 3D-consistent videos, establishing them as effective world simulators.
理想的 WFM 应展现出对物理规律的深刻理解,并生成遵循这些规律的未来观测。尽管我们的预训练 WFM 已具备一定程度的物理理解能力,并推动了该领域的发展,但仍然容易生成不符合物理规律的示例。我们认为,除了改进模型设计外,还需要在数据整理中增加步骤,以剔除物理上不合理的视频。虽然我们将强物理对齐的 WFM 留作未来工作,但我们仍然有兴趣衡量大规模数据驱动的预训练中自然涌现的直觉物理知识有多少。
An ideal WFM should exhibit a strong understanding of the laws of physics and produce future observations that respect them. While our pre-trained WFMs exhibit a certain level of physics understanding and advance the state-of-the-art, one can still easily generate examples that do not obey the laws of physics. We believe additional steps in data curation where physically implausible videos are removed are required, as well as improved model design. While we leave a strong physics-aligned WFM as future work, we are still interested in measuring how much intuitive physics naturally emerges from large-scale data-driven pre-training.
为了探索这一点,我们受 [ref] 启发,使用物理模拟引擎设计了一个受控基准数据集。我们生成基于物理的模拟,以测试预训练 WFM 对牛顿物理和刚体动力学的遵循程度。具体来说,我们使用模拟生成针对感兴趣物理定律的测试场景的物理正确的逼真视频。然后将这些参考“真实”视频与 WFM 在给定共享上下文(过去的观测和扰动)下生成的“预测”视频进行比较。
To explore this, we design a controlled benchmark dataset using a physics simulation engine, taking inspiration from [ref]. We generate physics-grounded simulations to test the adherence of our pre-trained WFMs to Newtonian physics and rigid body dynamics. Specifically, we use simulation to generate physically correct photorealistic videos of test scenarios specific to physical laws of interest. These reference “ground truth” videos are then compared with “predicted” videos produced by a WFM given shared context (past observations and perturbation).
合成数据生成。使用 PhysX 和 Isaac Sim,我们设计了八个 3D 场景,旨在评估不同的物理效应:
Synthetic data generation. Using PhysX and Isaac Sim, we design eight 3D scenarios aimed at evaluating different physical effects:
自由落体物体:物体掉落到平面上(重力、碰撞等)
Free-falling object(s): objects dropping on a plane (gravity, collision, etc.)
倾斜平面坡道:物体沿斜面滚下(重力、转动惯量等)
Tilted planar slope: objects rolling down an incline (gravity, moment of inertia, etc.)
U 形坡:物体沿 U 形坡滚下(势能、动能等)
U-shaped slope: objects rolling down a U-shaped slope (potential, kinetic energy, etc.)
稳定堆叠:处于平衡状态的物体堆叠(力平衡)
Stable stack: a stack of objects in equilibrium (balanced forces)
不稳定堆叠:处于不平衡状态的物体堆叠(重力、碰撞等)
Unstable stack: a stack of objects in imbalance (gravity, collision, etc.)
多米诺骨牌:矩形砖块依次倒下的序列(动量传递、碰撞等)
Dominoes: sequence of rectangular bricks falling in sequence (transfer of momentum, collision, etc.)
跷跷板:跷跷板两侧的物体(扭矩、转动惯量等)
Seesaw: objects on either side of a seesaw (torque, rotational inertia, etc.)
陀螺仪:平面上的旋转陀螺(角动量、进动等)
Gyroscope: a spinning top on a flat surface (angular momentum, precession, etc.)
对于每个场景,我们随机化动态物体的数量和类型(不同大小、纹理、形状),从 Omniverse 资源库中选取,同时随机化背景外观。我们模拟物体随时间的运动学状态,并从 4 个不同的静态相机视角渲染输出视频。总共渲染了 800 个 1080p 视频,每个视频 100 帧。每个模拟 rollout 中的物体位置确保从第一帧起全部可见,以避免存在性歧义。
For each scenario, we randomize the number and type of dynamic objects (varying sizes, textures, shapes), selecting from Omniverse assets, as well as the background appearance. We simulate the kinematic state of objects over time and render the output videos from 4 different static camera views. In total, we render 800 1080p videos of 100 frames in length. The objects in each simulation roll-out are positioned so that they are all visible from the first frame to avoid any existence ambiguity.
评估指标。我们通过将模拟的真实视频与 WFM 直接生成的输出进行比较,来评估其对物理规律的遵循程度。因此,为了生成未来观测,我们将 WFM 以真实视频的前几帧(1 帧或 9 帧)为条件。在适用时,我们还会额外以文本提示(通过专有 VLM 对条件帧进行描述获得)为条件,重点关注过去观测中模拟物体的运动学状态。请参阅图 20 了解模拟与预测场景的一些示例。评估使用以下指标:
Metrics. We are interested in assessing the adherence to physical laws by comparing the simulated ground-truth video to the output directly generated by the WFM. Therefore, to produce future observations, we condition our WFMs on the first few frames (either 1 or 9 frames) of the ground truth video. When applicable, we additionally condition a WFM on a text prompt (obtained using a proprietary VLM by captioning the conditioning frames), focusing on the kinematic state of the objects being simulated in the past observations. Please refer to fig. 20 for some examples of simulated versus predicted scenarios. For evaluation, we use the following metrics:
像素级指标。对于像素级比较,我们计算峰值信噪比(PSNR)和结构相似性指数(SSIM),以比较 WFM rollout 的预测帧与真实视频的参考帧。
Pixel-level metrics. For a pixel-level comparison, we compute the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) to compare a predicted frame from the WFM rollout with the reference frame from the ground truth video.
特征级指标。对于稍高层次的语义比较,我们计算预测帧与参考帧之间的 DreamSim 相似度分数,这是一种特征相似度指标。
Feature-level metrics. For a slightly higher-level semantic comparison, we calculate DreamSim similarity scores, a feature similarity metric, between the predicted and reference frames.
**对象级指标**。最后,由于我们最关心感兴趣对象如何受到持续物理现象的影响,我们使用跟踪来计算对象级指标,以消除混杂因素(背景变化、视觉质量等)。由于测试条件是合成生成的,我们可以获得场景中动态对象的真实实例分割掩码。我们使用 SAMURAI,将第一帧中的真实实例掩码传播到预测视频帧的其余部分以提取轨迹,从而量化对象级指标。我们计算每个帧和感兴趣对象的真实与预测对象掩码之间的交并比(IoU)。
Object-level metrics. Finally, since we care most about how objects of interest are impacted by the ongoing physical phenomenon, we use tracking to compute object-level metrics that eliminate confounders (background changes, visual quality, etc.). Since the test conditions are synthetically generated, we have access to the ground-truth instance segmentation masks of the dynamic objects in the scenes. Using SAMURAI, we propagate the ground-truth instance masks in the first frame through the rest of the predicted video frames to extract tracks, allowing us to quantify object-level metrics. We compute the intersection-over-union (IoU) between ground truth and predicted object masks for each frame and object of interest.
我们将这些指标在视频内的帧、评估集中的视频以及四个随机种子的 rollout 之间进行平均。PSNR 和 SSIM 在所有帧上计算,不包括用于条件生成的帧。
We average these metrics across frames in a video, across videos in the evaluation set, and across four random seeds for rollouts. PSNR and SSIM are computed on all frames, excluding the ones used for conditioning.
**结果**。物理对齐的定量结果在表 20 中概述。基于定量和定性结果,我们做出以下观察。不出所料,模型在更多帧作为条件输入时能够更好地预测整体对象运动学(这使我们能够更好地推断一阶和二阶量,如速度和加速度)。
Results. Quantitative results on physical alignment are outlined in Table 20. Based on quantitative and qualitative results, we make the following observations. Unsurprisingly, the models are able to better predict the overall object kinematics with more frames as conditioning input (which allows us to better infer 1st and 2nd order quantities such as speed and acceleration).
从表中我们还发现,在 9 帧条件设置下,我们的扩散 WFM 在像素级预测上优于自回归 WFM。这与我们的视觉观察一致,即基于扩散的 WFM 生成的视频具有更高的视觉质量。我们还注意到,我们的结果并不表明更大的模型在物理对齐上表现更好。虽然我们观察到更大的模型生成的视频具有更高的视觉质量,但所有 WFM 在物理遵循方面同样困难,需要更好的数据策划和模型设计。
From the table, we also find that our diffusion WFMs perform better in pixel-level prediction than our autoregressive WFMs on the 9-frame conditional setting. This correlates with our visual observation that the diffusion-based WFMs render videos with higher visual quality. We also note that our results do not suggest that the larger model performs better on our physics alignment. While we observe larger models render videos with higher visual quality, all the WFMs equally struggle with physics adherence and require better data curation and model design.
更一般地,我们观察到上述刚体模拟已经测试了我们 WFM 的极限,作为识别特定失败案例的有价值工具。这些失败案例从低级问题(如对象不持久(对象的自发出现和消失)和变形(形状变化))到更复杂的问题(如不可信的运动学、违反重力等)。我们相信这种结构化模拟提供了一种有用的方法来测试物理对齐。因此,我们打算随着时间的推移改进它们,纳入更复杂的场景,增强逼真度以弥合模拟与真实之间的差距(因为 WFM 预训练数据由真实视频组成),并完善我们的评估指标,以更全面地评估物理理解。
More generally, we observe that the rigid-body simulations described above already test the limits of our WFMs, serving as valuable tools for identifying specific failure cases. These range from low-level issues like object impermanence (spontaneous appearance and disappearance of objects) and deformation (shape changes) to more complex problems such as implausible kinematics, violation of gravity, etc. We believe such structured simulations offer a useful methodology to test physics alignment. We, therefore, intend to improve them over time by incorporating more complex scenarios, enhancing photorealism to bridge the sim-to-real gap (since WFM pre-training data consists of real videos), and refining our evaluation metrics for a more comprehensive assessment of physical understanding.
在本节中,我们展示了如何对 Cosmos WFM 进行微调以支持各种物理 AI 应用。我们提供的示例包括:通过相机控制对 WFM 进行后训练,实现可导航的 3D 视觉世界生成;在两种不同的机器人设置上,针对两种不同的机器人操作任务,通过动作控制对 WFM 进行后训练;以及通过多视角支持对 WFM 进行后训练,用于训练自动驾驶智能体。
In this section, we demonstrate how our Cosmos WFMs can be fine-tuned to support diverse Physical AI applications. We include examples from post-training our WFM with camera control to achieve 3D navigable visual world generation, post-training our WFM with action control on two different robotic setups for two different robotic manipulation tasks, and post-training our WFM with multi-view support for training autonomous driving agents.
通过相机姿态条件化,我们将相机控制集成到 Cosmos-Predict1-7B-Video2World 中,使其成为一个有效的 3D 世界模拟器。我们将后训练的世界基础模型称为 Cosmos-Predict1-7B-Video2World-Sample-CameraCond。我们专注于从单个参考输入图像生成 3D 世界,利用相机控制从指定的相机轨迹生成时间上连贯且 3D 一致的视频模拟,其中视角的变化与场景的底层 3D 结构对齐。
Through camera pose conditioning, we integrate camera control into Cosmos-Predict1-7B-Video2World, making it an effective 3D world simulator. We term the result post-trained WFM as Cosmos-Predict1-7B-Video2World-Sample-CameraCond. We focus on generating 3D worlds from a single reference input image, leveraging camera control to produce temporally coherent and 3D-consistent video simulations from the specified camera trajectories, where changes in perspective align with the underlying 3D structure of the scene.
我们使用 DL3DV-10K,这是一个大规模静态场景视频数据集,用于此任务。作为预处理步骤,我们将所有视频切分为包含 256 帧的片段。为了获得片段内所有帧的密集相机位姿标注,我们使用 GLOMAP 对切分后的片段运行运动恢复结构(structure-from-motion)。我们将第一帧的相机位姿设为单位变换,并计算所有后续帧的相对相机位姿。我们还使用专有的视觉语言模型(VLM)为视频生成描述,以获得将视频描述为静态场景的文本提示。
We use DL3DV-10K, a large-scale video dataset of static scenes, for this task. As a preprocessing step, we chunk all videos into clips with 256 frames. To obtain dense camera pose annotations for all frames within a clip, we run structure-from-motion on the chunked clips using GLOMAP. We set the camera pose of the first frame to be the identity transform and compute the relative camera poses for all subsequent frames. We also use a proprietary VLM to caption the videos to obtain text prompts that describe the videos as static scenes.
我们通过将采样的潜在嵌入与普吕克嵌入拼接来添加相机控制条件,普吕克嵌入与潜在嵌入具有相同的空间维度。具体来说,给定相机姿态,我们通过以下方式计算普吕克坐标:
We add camera control conditioning by concatenating the sampled latent embeddings with Plücker embeddings, which have the same spatial dimensions as the latent embeddings. Specifically, given the camera pose, we compute the Plücker coordinates via
其中 \(\mathbf{c}\) 是相机中心位置,\(\mathbf{d}\) 是每个潜在像素的单位射线方向(将潜在嵌入视为下采样图像)。所有相机姿态均相对于初始帧。Cosmos-Predict1-7B-Video2World 模型使用的 Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p 具有 \(8\times\) 的时间压缩率,因此对于每 8 帧,我们使用第 4 帧的普吕克嵌入与相应的潜在表示进行拼接。
where \(\mathbf{c}\) is the camera center location and \(\mathbf{d}\) is the unit ray direction of each latent pixel (where the latent embedding is treated as a downsampled image). All the camera poses are relative with respect to the initial frame. The Cosmos-Tokenize1-CV8 \(\times\) 8 \(\times\) 8-720p used by Cosmos-Predict1-7B-Video2World models has a temporal compression rate of \(8\times\) , and thus for every 8 frames, we use the Plücker embedding at the 4th frame to concatenate with the corresponding latent representation.
我们将训练视频的输入帧调整为 \(704\times 1252\),并使用反射填充至 \(704\times 1280\)。训练期间采样 57 帧。训练目标和其他超参数与基础 Diffusion WFM 训练(第 5.1.3 节)相同。
We resized the input frames of our training videos to \(704\times 1252\) and padded them to \(704\times 1280\) with reflection. We sample 57 frames during training. The training objective and other hyper-parameters are the same as the base Diffusion WFM training (section 5.1.3).
我们假设给定一张世界参考图像,并从输入图像生成未来展开的视频。我们与 CamCo 进行比较,CamCo 是在此设置下用于相机可控视频生成的最先进模型。为了公平比较,我们使用了也在 DL3DV-10K 训练集上微调过的 CamCo 模型。由于我们的后训练 WFM 生成 57 帧,而 CamCo 只能生成 14 帧,我们比较相同的 57 帧轨迹,其中对 CamCo 进行时间下采样 \(4\times\)。CamCo 的视频分辨率限制为 \(256\times 256\)。此外,我们对输入图像和测试帧进行最大限度的中心裁剪以进行评估。
We assume a single reference image of the world is given and generate the future rollout as a video from the input image. We compare against CamCo, the state-of-the-art model for camera-controllable video generation under this setup. For a fair comparison, we use the CamCo model that was also fine-tuned on the DL3DV-10K training set. As our post-trained WFM generates 57 frames and CamCo can only generate 14 frames, we compare the same 57-frame trajectories where we temporally downsample by \(4\times\) for CamCo. The video resolution from CamCo is limited to \(256\times 256\). We additionally maximally center-crop the input image and test frames for evaluation.
对于测试数据,我们使用第 5.3.1 节中描述的 RealEstate10K 测试集中的相同 500 个样本。我们使用初始帧作为参考图像,并使用数据集提供的相机轨迹作为相机控制输入,此外我们还对其进行重新缩放,使得轨迹两端之间的距离归一化为 1。
For the test data, we use the same 500 samples from the RealEstate10K test set previously described in section 5.3.1. We use the initial frame as the reference image and camera trajectories provided by the dataset as the camera control input, which we additionally rescale such that the distance between two ends of trajectories is normalized to 1.
**指标**。遵循 ,我们从两个方面评估后训练世界模型的相机可控性:视频生成质量和 3D 一致性。对于视频质量,我们使用 Fréchet Inception Distance(FID)和 Fréchet Video Distance(FVD)分别评估帧级和视频级质量。我们使用与参考视频相同的测试数据来计算指标(注意它们不用于像素级比较)。
Metrics. Following , we evaluate the camera controllability of the post-trained world model in two aspects: video generation quality and 3D consistency. For video quality, we use the Fréchet Inception Distance (FID) and the Fréchet Video Distance (FVD) to assess the qualities at the frame and video levels, respectively. We use the same test data as the reference videos to compute the metrics (note that they are not used for pixel-level comparisons).
对于 3D 一致性,我们通过运动恢复结构库重新估计相机位姿的能力来评估,并将结果与输入相机控制轨迹进行比较。给定视频中的 \(N\) 帧,我们将相机轨迹误差量化为两个项:平均旋转误差 \(\epsilon_{\text{rot}}\) 和平移误差 \(\epsilon_{\text{trans}}\),分别定义为
For 3D consistency, we evaluate via the ability of structure-from-motion libraries to re-estimate the camera poses, and we compare the results against the input camera control trajectories. Given \(N\) frames in the video, we quantify the camera trajectory error into two terms: the average rotation error \(\epsilon_{\text{rot}}\) and translation error \(\epsilon_{\text{trans}}\), defined respectively as
其中 \(\mathbf{R}_{i}\) 和 \(\mathbf{t}_{i}\) 是第 \(i\) 帧的输入旋转和平移(作为地面真值),\(\mathbf{\hat{R}}_{i}\) 和 \(\mathbf{\hat{t}}_{i}\) 是重新估计的量。为了考虑相机位姿估计结果中相似变换带来的歧义,我们遵循 并对预测的相机轨迹进行 Procrustes 分析以与地面真值对齐。
where \(\mathbf{R}_{i}\) and \(\mathbf{t}_{i}\) are the input rotation and translation of the \(i\)-th frame (serving as ground truth), and \(\mathbf{\hat{R}}_{i}\) and \(\mathbf{\hat{t}}_{i}\) are the re-estimated quantities. To account for ambiguities from camera pose estimation results up to a similarity transformation, we follow and run Procrustes analysis on the predicted camera trajectories to align against the ground truth.
**对比。** 我们在表 22 中展示了结果。首先,我们后训练的 WFM 能够生成逼真且连贯的 3D 世界。这体现在较低的 FID/FVD 分数(更高的视觉质量)和较高的相机姿态估计成功率上。Cosmos-Predict1-7B-Video2World-Sample-CameraCond 展示了更好的相机控制能力,因为其相机轨迹重估结果与原始控制输入显著更接近。
Comparisons. We present the results in Table 22. First, our post-trained WFM can generate realistic and coherent 3D worlds. This is evidenced by the lower FID/FVD scores (higher visual quality) and the higher camera pose estimation success rate. Cosmos-Predict1-7B-Video2World-Sample-CameraCond demonstrates better camera control, as the camera trajectory re-estimation is significantly closer to the original control input.
**视觉对比。** 我们在图 21 中提供了视觉对比。虽然 CamCo 难以生成输入图像之外的内容,但 Cosmos-Predict1-7B-Video2World-Sample-CameraCond 能够有效生成符合 3D 世界结构的视觉内容。请注意,两个模型都在 DL3DV-10K 上进行了后训练,并在 RealEstate10K 数据集上进行了评估,这引入了训练与测试之间的显著分布偏移。Cosmos 模型成功克服了这种分布偏移,同时展示了其对未见过的输入相机轨迹的泛化能力。
We also provide visual comparisons in Fig. 21. While CamCo struggles to generate content beyond the input image, Cosmos-Predict1-7B-Video2World-Sample-CameraCond effectively generates visuals that adhere to the structure of a 3D world. Note that both models were post-trained on DL3DV-10K and evaluated on the RealEstate10K dataset, which introduces a significant distribution shift between training and testing. The Cosmos model successfully overcomes this distribution shift while also demonstrating its capability to generalize to unseen input camera trajectories.
**定性结果。** 图 22 展示了我们使用类似摇杆的相机控制输入(包括前进、后退、左转和右转)所得到的结果。这演示了用户可以通过摇杆控制模型生成未来视频帧,从而在模拟世界中导航的应用场景。物理 AI 智能体也可以利用这种控制来预测不同场景下的世界未来状态。
Qualitative results. Fig. 22 shows our results from joystick-like control input on the camera, including moving forward, moving backward, rotating left, and rotating right. This demonstrates the use case where one can navigate the simulated world using a joystick to control the model in generating future video frames. A Physical AI agent could also use such control to predict the future of the world under different scenarios.
**多样性展示。** 为了展示生成的多样性,我们在图 23 中展示了使用相同输入图像和相机控制但不同随机种子时的生成结果。Cosmos-Predict1-7B-Video2World-Sample-CameraCond 能够生成不同的世界,同时保持视频中的 3D 空间和时间连贯性。这可用于根据当前状态模拟多种可能的未来。
To show the diversity of the generation, we show generation results from the same input image and camera control with different random seeds in Fig. 23. Cosmos-Predict1-7B-Video2World-Sample-CameraCond is able to generate different worlds while still maintaining 3D spatial and temporal coherence in the videos. This could be used to simulate different possible futures given the current states.
世界模型有潜力成为机器人操作的强大规划器和模拟器。在此,我们展示了如何针对两个任务微调预训练的世界模型(WFM):(1)基于指令的视频预测和(2)基于动作的下一帧生成。对于基于指令的视频预测,输入是机器人的当前视频帧以及文本指令,输出是机器人遵循指令的预测视频。对于基于动作的下一帧预测,输入是机器人的当前视频帧以及当前帧与下一帧之间的动作向量,输出是预测的下一帧,显示机器人执行指定动作的结果。给定一系列动作,模型可以自回归运行,预测机器人执行给定动作的视频。
A world model has the potential to serve as a powerful planner and simulator for robotic manipulation. Here, we demonstrate how we fine-tune our pre-trained WFMs for two tasks: (1) instruction-based video prediction and (2) action-based next-frame generation. For instruction-based video prediction, the input is the current video frame of a robot as well as a text instruction, and the output is a predicted video of the robot following the instruction. For action-based next-frame prediction, the input is the current video frame of a robot as well as an action vector between the current and next frame, and the output is the predicted next frame showing the result of the robot performing the specified action. Given a sequence of actions, the model can be run autoregressively to predict a video of the robot executing the given actions.
我们为上述两个任务整理了两个数据集。对于基于指令的视频预测,我们创建了一个内部数据集,称为 Cosmos-1X 数据集。该数据集包含约 200 小时的以自我为中心的视频,由 1x.Tech 的人形机器人 EVE 执行各种任务时拍摄,包括导航、叠衣服、清洁桌子、拾取物体等。从原始视频中,我们选取了约 12,000 个片段,时长从 1 到 9 秒不等。每个片段都标注有一句话指令,之后使用专有的 VLM 进行上采样。视频以 30 FPS 的帧率、512×512 的分辨率拍摄。
We curate two datasets for the two tasks described above. For instruction-based video prediction, we created an internal dataset called the Cosmos-1X dataset. It comprises approximately 200 hours of egocentric videos captured by EVE, a humanoid robot from 1x.Tech performing a variety of tasks, including navigation, folding clothes, cleaning tables, picking up objects, etc. From the raw videos, we selected approximately 12,000 episodes ranging from 1 to 9 seconds. Each episode is labeled with a one-sentence instruction, which is later upsampled with a proprietary VLM. The videos are captured at 30 FPS with a resolution of 512×512.
对于基于动作的下一帧生成,我们使用了公开数据集 Bridge,其配置与先前工作相同,以便进行比较。Bridge 数据集包含约 20,000 个片段,展示机器人手臂在厨房环境中执行不同任务的第三人称视角,视频分辨率为 320×256,帧率为 5 FPS。对于每个视频帧,相应的动作定义为夹爪坐标系中的 7 维向量 (Δx, Δy, Δz, Δθr, Δθp, Δθy, ΔGripper),与 OpenVLA 中的定义一致。
For action-based next-frame generation, we used a public dataset called Bridge, with the same configuration as a prior work for comparison. The Bridge dataset includes approximately 20,000 episodes of third-person views of a robot arm performing different tasks in a kitchen environment, with videos of 320×256 resolution captured at 5 FPS. For each video frame, the corresponding action is defined as a 7-dimensional vector in the gripper coordinate space (Δx, Δy, Δz, Δθr, Δθp, Δθy, ΔGripper) as in OpenVLA.
我们对 Cosmos-Predict1-7B-Video2World(第 5.1 节)和 Cosmos-Predict1-5B-Video2World(第 5.2 节)进行了微调,以完成基于指令的视频预测和基于动作的下一帧预测任务。
We fine-tune both our Cosmos-Predict1-7B-Video2World (section 5.1) and Cosmos-Predict1-5B-Video2World (section 5.2) for instruction-based video prediction and action-based next-frame prediction tasks.
对于基于指令的视频预测,我们在基础 WFM 之上构建了两个模型。第一个称为 Cosmos-Predict1-7B-Video2World-Sample-Instruction,第二个称为 Cosmos-Predict1-5B-Video2World-Sample-Instruction。我们计算指令的 T5 嵌入,并通过交叉注意力将其添加到基础模型的微调中。
For instruction-based video prediction, we build two models based on the base WFMs. The first is called Cosmos-Predict1-7B-Video2World-Sample-Instruction, and the second is called Cosmos-Predict1-5B-Video2World-Sample-Instruction. We compute the T5 embedding of the instruction, which is added to the fine-tuning of the base model via cross-attention.
对于基于动作的下一帧预测,我们同样在基础 WFM 之上构建了两个模型。第一个称为 Cosmos-Predict1-7B-Video2World-Sample-ActionCond,第二个称为 Cosmos-Predict1-5B-Video2World-Sample-ActionCond。
For action-based next-frame prediction, we also build two models based on the base WFMs. The first one is called Cosmos-Predict1-7B-Video2World-Sample-ActionCond, and the second one is called Cosmos-Predict1-5B-Video2World-Sample-ActionCond.
由于动作是一种在预训练期间未遇到的新模态,我们在模型中引入了额外的模块用于条件化。对于 Cosmos-Predict1-5B-Video2World-Sample-ActionCond,我们添加了一个动作嵌入器 MLP,将动作向量投影为张量,然后通过交叉注意力将其并入模型。对于 Cosmos-Predict1-7B-Video2World-Sample-ActionCond,我们也添加了一个动作嵌入器 MLP,将动作预测为张量,但不同的是,我们将其添加到 DiT 模块的时间戳嵌入中,从而并入模型。
Since action is a new modality not encountered during pre-training, we introduce additional modules inside our models for conditioning. For Cosmos-Predict1-5B-Video2World-Sample-ActionCond, we add an action embedder MLP to project the action vector into a tensor, which is then incorporated into the model via cross-attention. For Cosmos-Predict1-7B-Video2World-Sample-ActionCond, we also add an action embedder MLP to predict the action into a tensor but instead, incorporate it into the model by adding it to the timestamp embedding of the DiT modules.
对于基于指令的视频预测,我们在 Cosmos-1X 数据集上对 VideoLDM 进行微调,并得到 VideoLDM-Instruction 作为比较的基线。为了评估模型的视频生成性能,我们定义了以下维度:
For instruction-based video prediction, we fine-tune VideoLDM on the Cosmos-1X dataset and obtain VideoLDM-Instruction as a baseline for comparison. To evaluate the video generation performance of the models, we define the following dimensions:
指令遵循:生成的视频是否与输入的语言指令对齐?
Instruction following: Is the generated video aligned with the input language instruction?
物体恒存:场景中存在的物体是否在生成的视频中始终存在?
Object permanence: Do objects present in the scene remain throughout the generated video?
真实性:生成的视频是否忠实地表示真实世界,没有意外的虚构物体?
Verity: Does the generated video faithfully represent the real world without unexpected imaginary objects?
总体:生成的视频是否合理,足以让机器人据此进行规划?
Overall: Is the generated video reasonable for the robot to plan accordingly?
人类评估员的任务是观察由不同模型生成但具有相同语言指令的一对匿名视频,并按照上述维度进行比较。十名人类评估员对 23 个测试片段进行了评估。统计结果总结在图 24 中。
Human evaluators are tasked to observe a pair of anonymous videos generated by different models but with the same language instruction and compare them along the dimensions listed above. A group of ten human evaluators performed the evaluation over 23 test episodes. The statistical results are summarized in Fig. 24.
如图所示,我们发现 Cosmos-Predict1-7B-Video2World-Sample-Instruction 和 Cosmos-Predict1-5B-Video2World-Sample-Instruction 在四个评估维度上均优于 VideoLDM-Instruction。Cosmos-Predict1-7B-Video2World-Sample-Instruction 获得了 78.3%的总体偏好,而 VideoLDM-Instruction 为 13.0%。Cosmos-Predict1-5B-Video2World-Sample-Instruction 也取得了优于基于扩散的 VideoLDM-Instruction 的性能。两个微调后的 WFM 的一些预测视频帧如图 25 所示,展示了预测视频的质量。
As shown, we find that both Cosmos-Predict1-7B-Video2World-Sample-Instruction and Cosmos-Predict1-5B-Video2World-Sample-Instruction perform better than VideoLDM-Instruction along the four evaluation dimensions. Cosmos-Predict1-7B-Video2World-Sample-Instruction achieved \(78.3\%\) overall preference compared to \(13.0\%\) for VideoLDM-Instruction. Cosmos-Predict1-5B-Video2World-Sample-Instruction has also achieved better performance than diffusion-based VideoLDM-Instruction. Some predicted video frames for both fine-tuned WFMs are presented in Fig. 25, which shows the quality of the predicted videos.
对于基于动作的下一帧预测,我们在 Bridge 数据集上微调了模型。作为基线,我们微调 IRASim 以得到基于动作的下一帧预测模型 IRASim-Action。我们自回归地进行下一帧预测以生成视频。为了评估视频生成质量,我们将生成的视频与从官方 Bridge 测试集中随机选择的 100 个片段的地面真实视频进行比较。
For action-based next frame prediction, we fine-tuned our models on the Bridge dataset. As a baseline, we fine-tune IRASim to derive an action-based next-frame prediction model IRASim-Action. We perform the next-frame prediction autoregressively to generate videos. To evaluate video generation quality, we compare the generated videos against ground truth videos over 100 episodes randomly selected from the official Bridge test set.
计算得到的指标总结在表 23 中,包括 PSNR、SSIM、Latent L2 和 FVD。如图所示,Cosmos-Predict1-5B-Video2World-Sample-ActionCond 和 Cosmos-Predict1-7B-Video2World-Sample-ActionCond 模型均优于基线模型(IRASim-Action)。一些预测视频帧如图 26 所示,展示了预测视频与地面真实视频相比的质量。
The computed metrics are summarized in Table 23, including PSNR, SSIM, Latent L2, and FVD. As shown, both Cosmos-Predict1-5B-Video2World-Sample-ActionCond and Cosmos-Predict1-7B-Video2World-Sample-ActionCond models outperform the baseline model (IRASim-Action). Some predicted video frames are presented in Fig. 26, which shows the quality of the predicted videos compared to the ground truth.
面向野外驾驶场景的世界模型有潜力成为训练自动驾驶智能体的强大仿真引擎。由于大多数自动驾驶车辆配备了多个朝向不同方向的摄像头,理想的自动驾驶世界模型也应该是多视角的,最好与目标车辆中传感器的精确配置相匹配。在此,我们展示了如何微调我们预训练的世界基础模型(WFM),以创建用于自动驾驶任务的多视角世界模型。
A world model for in-the-wild driving scenes has the potential to serve as a powerful simulation engine for training autonomous driving agents. As most autonomous vehicles are equipped with multiple cameras viewing different directions, an ideal world model for an autonomous vehicle should also be a multi-view one, preferably matching the precise setup of the sensors in the target vehicle. Here, we demonstrate how we fine-tune our pre-trained WFM to create a multi-view world model for autonomous driving tasks.
我们整理了一个内部数据集,称为真实驾驶场景(RDS)数据集。该数据集包含约 360 万个 20 秒环视视频片段(相当于约 20,000 小时的数据),这些片段使用 NVIDIA 内部驾驶平台捕获。每个片段由六个摄像头视角记录:前、左、右、后、左后和右后。此外,数据集还包括我们用于构建轨迹数据的自我运动信息。我们使用前摄像头视频的记录时间戳来同步所有其他视角的帧。
We curate an internal dataset called the Real Driving Scene (RDS) dataset. It comprises approximately 3.6 million 20-second surround-view video clips (equivalent to approximately 20,000 hours of data) captured using an NVIDIA internal driving platform. Each clip is recorded from six camera views: front, left, right, rear, rear-left, and rear-right. In addition, the dataset includes ego-motion information that we use to construct the trajectory data. We use the recorded timestamps of the front camera video to synchronize the frames of all other views.
该数据集是从一个大型标注数据语料库中选取的,以匹配目标的数据属性分布。具体的属性标签包括:
This dataset was selected from a large labeled data corpus to match a target distribution of data attributes. The specific attribute tags include:
周围车辆密度(例如:无、低、中、高)
Contender vehicle density (e.g., none, low, medium, high)
天气(例如:晴朗、下雨、下雪、雾)
Weather (e.g., clear, raining, snowing, fog)
光照(例如:白天、夜晚)
Illumination (e.g., day, night)
自车速度(例如,静止、低速、本地、高速)
Ego vehicle speed (e.g., standing, low, local, highway speeds)
自车行为(例如,高、中、低曲率轨迹和加速度)
Ego vehicle behavior (e.g., high, medium, low curvature trajectories and accelerations)
道路类型/人口密度(基于 OpenStreetMap 定义:乡村、住宅区、城市)。
Road type/population density (based on OpenStreetMap definitions: rural, residential, urban).
此外,通过第二次数据挖掘运行对数据集进行了扩充,以确保包含稀有道路结构(例如,收费站、桥梁、隧道、减速带等)的片段达到最小数量。最后,每个摄像头视角的视频分别添加字幕,以模板文本字符串开头:“视频由安装在汽车上的摄像头拍摄。摄像头朝向 前方|左方|右方|后方|左后方|右后方。”
Additionally, the dataset was augmented through a second data-mining run to ensure a minimum number of clips containing rare road structures (e.g., tollbooths, bridges, tunnels, speed bumps, etc.). Finally, videos from each camera view are captioned separately, starting with a template text string: “The video is captured from a camera mounted on a car. The camera is facing forward|left|right|backward|rear-left|rear-right.”
我们使用 RDS 数据集将 Cosmos-Predict1-7B-Text2World(第 5.1 节)微调为多视角世界模型。为确保多视角下视频生成的一致性,我们略微修改了第 5.1 节所述的架构设计,并微调 WFM 以同时从六个摄像头生成视频。
We fine-tune our Cosmos-Predict1-7B-Text2World (section 5.1) into a multiple-view world model using the RDS dataset. To ensure consistent video generation across multiple views, we slightly modify the architectural design described in section 5.1 and fine-tune the WFM to generate videos from all six cameras simultaneously.
我们构建了三个多视角世界模型,总结于表 21。第一个模型名为 Cosmos-Predict1-7B-Text2World-Sample-MultiView,它是一个多视角世界模型,可根据文本提示输入生成六个摄像头视角。第二个模型名为 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond,该模型基于 Cosmos-Predict1-7B-Text2World-Sample-AV-MultiView 构建,并额外将轨迹输入作为条件输入信号。最后一个模型 Cosmos-Predict1-7B-Video2World-Sample-MultiView 从 Diffusion-7B-Video2World-Sample-MultiView 模型微调而来,以支持基于视频的条件生成,其实现方式是将先前帧纳入生成过程。Cosmos-Predict1-7B-Video2World-Sample-MultiView 可以接收 Cosmos-Predict1-7B-Text2World-Sample-MultiView 的视频输出并生成其扩展。所有三个模型均输出 6 个视角、57 帧、分辨率为 \(848\times 480\) 的视频。
We build three multi-view world models, summarized in table 21. The first one is called Cosmos-Predict1-7B-Text2World-Sample-MultiView, which is a multi-view world model that can generate six camera views based on a text prompt input. The second one is called Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond. This model is built on top of Cosmos-Predict1-7B-Text2World-Sample-AV-MultiView and takes an additional trajectory input as the conditional input signal. The final model, Cosmos-Predict1-7B-Video2World-Sample-MultiView, is fine-tuned from the Diffusion-7B-Video2World-Sample-MultiView model to support video-based conditioning. It achieves this by incorporating previous frames into the generation process. Cosmos-Predict1-7B-Video2World-Sample-MultiView can take the video output from Cosmos-Predict1-7B-Text2World-Sample-MultiView and generate its extension. All three models output 6 views of 57 frames of video at a resolution of \(848\times 480\) .
视角无关的位置嵌入和视角嵌入。我们没有将基于 FPS 的 3D RoPE 位置嵌入扩展以包含额外的视角维度,而是选择对每个视角独立使用第 5.1 节所述的相同位置嵌入。为了表示视角差异,我们修改去噪函数 \(D_{\theta}\),使其接受额外的视角嵌入作为输入。也就是说,相机视角信息通过全局视角嵌入而非位置嵌入提供。
View-independent positional embedding and view embedding. Instead of extending the FPS-aware 3D RoPE Positional Embedding to include an additional view dimension, we opt to use the same positional embedding described in section 5.1 independently to each view. To represent view differences, we modify the denoising function \(D_{\theta}\) to take an additional view embedding as input. That is, the camera view information is supplied through global view embeddings instead of positional embedding.
视角相关的交叉注意力。在我们的多视角设置中,同一场景的六个视角各自具有不同的视频描述。虽然我们将六个视角整体视为扩散过程的状态,并在六个视角的所有元素之间执行自注意力以进行去噪,但我们发现对文本输入采用视角相关的交叉注意力是有益的。具体来说,每个视角的交叉注意力操作仅关注该特定视角的文本描述。请注意,在我们的数据集中,每个视角具有不同的视频描述。通过视角嵌入和视角相关的交叉注意力,我们从微调 Cosmos-Predict1-7B-Text2World 得到 Cosmos-Predict1-7B-Text2World-Sample-MultiView。
View-dependent cross-attention. In our multi-view setting, each of the six views of the same scene would have a different video description. While we treat all six views as a whole as the state of the diffusion process and perform self-attention among all the elements in the six views for denoising, we find it beneficial to employ view-dependent cross-attention for textual inputs. Specifically, the cross-attention operation for each view only attends to the textual description for the specific view. Note that each view has a different video description in our dataset. With the view embedding and view-dependent cross-attention, we derive Cosmos-Predict1-7B-Text2World-Sample-MultiView from fine-tuning Cosmos-Predict1-7B-Text2World.
轨迹控制条件。可选地,除了文本条件之外,我们微调模型以生成符合给定未来轨迹路径的视频,从而实现对智能体更精确的控制。这使得能够生成独特的驾驶场景,既符合真实世界数据记录的驾驶轨迹,又符合输入文本描述所指定的驾驶环境。微调后的模型为 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond。
Trajectory control condition. Optionally, in addition to the text condition, we fine-tune the model to produce videos that conform to the given future trajectory paths to enable more precise control of the agent. This enables the generation of unique driving scenarios that adhere to both driving trajectories recorded by real-world data and driving environments specified by the input text descriptions. The fine-tuned model is Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond.
我们将轨迹定义为三维空间中的 64 个点的序列,表示智能体从初始位置 \((0,0,0)\) 到最终目的地的平移序列,每个点之间间隔 0.1 秒。我们计算轨迹输入的嵌入,并将结果作为微调后的 Cosmos-Predict1-7B-Video2World 模型去噪器的条件输入。我们注意到,通过提供每个时间间隔的动作向量,可以遵循先前的工作或如机器人操作任务(第 6.2 节)那样,实现更细粒度的控制信号。我们将此类扩展留待未来工作。
We define a trajectory as a sequence of 64 points in 3D space, representing a sequence of translations of the agent from the initial position \((0,0,0)\) to the final destination, with each point separated by a 0.1-second interval. We compute the embedding of the trajectory input and make the result a conditional input to the denoiser of the fine-tuned Cosmos-Predict1-7B-Video2World model. We note that it is possible to achieve more fine-grained control signals by giving a per-interval action vector, following prior works or as in the robotic manipulation task (section 6.2). We leave such extensions for future work.
我们首先在图 27 中展示基于文本条件的定性结果。使用 Cosmos-Predict1-7B-Text2World-Sample-MultiView,我们生成了一个包含六个视角的 57 帧视频,然后使用 Cosmos-Predict1-7B-Video2World-Sample-MultiView 模型将其扩展到 201 帧。在图 28 中,我们展示了预训练的世界模型如何增强泛化能力,从而能够生成 RDS 数据集中罕见或超出分布的场景,例如在河上行驶。最后,图 29 展示了 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 的结果,其中自车精确地遵循输入轨迹。
We first present text-conditioned qualitative results in Fig. 27. Using Cosmos-Predict1-7B-Text2World-Sample-MultiView, we generate a 57-frame video with six views, which is then extended to 201 frames using the Cosmos-Predict1-7B-Video2World-Sample-MultiView model. In Fig. 28, we demonstrate how the pre-trained world model enhances generalization, enabling the generation of rare or out-of-domain scenes from the RDS dataset, such as driving on a river. Lastly, Fig. 29 showcases the results from Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond, where the ego car accurately follows the input trajectory.
对于定量结果,作为基线,我们采用相同的微调方案对 VideoLDM 进行微调,得到名为 VideoLDM-MultiView 的多视角世界模型。我们使用一组评估指标来衡量视频生成质量、多视角一致性和轨迹跟随精度。为了评估视频生成质量,我们使用 1000 个样本计算得分。对于一致性相关指标,为了更好地理解不同模型在不同场景下的行为,我们将真实轨迹分为四类:直行、左转、右转和其他(包括静态或复杂运动)。对于每个类别,我们收集了 200 个样本及其对应的提示和条件,总计 800 个样本。下面,我们提供指标和结果的详细描述。
For quantitative results, as a baseline, we followed the same fine-tuning recipe to fine-tune VideoLDM to derive a multi-view world model called VideoLDM-MultiView. We use a set of evaluation metrics measuring video generation quality, multi-view consistency, and trajectory following accuracy. To evaluate video generation quality, we use 1000 samples to compute the scores. For consistency-related metrics, to better understand different models’ behaviors under different scenarios, we categorize the ground-truth trajectories into four types: moving forward, turning left, turning right, and others (including static or complex movements). For each category, we gathered 200 samples and their corresponding prompts and conditions, totaling 800 samples. Below, we provide detailed descriptions of the metrics and the results.
生成质量。我们使用 Fréchet Inception Distance(FID)和 Fréchet Video Distance(FVD)来衡量生成视频相对于真实视频的质量。我们首先从每个视频中提取 16 帧,计算每个视角的得分,然后报告每种方法在所有视角上的平均得分。如表 24 所示,我们发现 Cosmos-Predict1-7B-Text2World-Sample-MultiView 和 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 在这两个指标上都显著优于 VideoLDM-MultiView,证明了我们预训练的基于扩散的 7B WFM 相对于 VideoLDM-MultiView 基线的优越质量。
Generation quality. We utilize Fréchet Inception Distance (FID) and Fréchet Video Distance (FVD) to measure the quality of the generated videos relative to the real ones. We first calculate a score per view by extracting 16 frames from each video. We then report the average score across all views per method. As shown in Table 24, we find both Cosmos-Predict1-7B-Text2World-Sample-MultiView and Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond significantly outperform VideoLDM-MultiView on both metrics, demonstrating the superior quality of our pre-trained 7B Diffusion-based WFM over the VideoLDM-MultiView baseline.
多视角一致性。我们使用第 5.3.1 节中公式化的 Sampson 误差的扩展版本来量化生成的多视角视频的几何一致性。由于 RDS 数据集中的真实视频共享相似的鱼眼相机内参,我们使用中值标定将关键点去畸变到均匀大小为 \(960\times 540\) 且水平视场角为 120 度的常规针孔相机。在此设置下,为生成的多视角视频计算两个指标:
Multi-view consistency. We use an extended version of the Sampson error formulated in Section 5.3.1 to quantify the geometry consistency of the generated multi-view videos. As the ground-truth videos in our RDS dataset share similar fisheye camera intrinsic parameters, we use the median calibration to undistort the keypoints to a regular pinhole camera with a uniform size of \(960\times 540\) and 120 degrees of horizontal FoV. Under this setting, two metrics are computed for the generated multi-view video:
时间 Sampson 误差(TSE)衡量每个相机生成的内容在时间上是否一致。它是每个视角相邻帧的 Sampson 误差的中位数。
Temporal Sampson Error (TSE) measures whether the content generated for each camera is consistent over time. It is the median Sampson error of adjacent frames for each of the views.
跨视角萨姆森误差(CSE)衡量多视角一致性是否随时间保持。它是不同生成视角之间萨姆森误差的时间平均。CSE 中使用的基础矩阵是通过累积所有时间帧的关键点估计得到的。
Cross-view Sampson Error (CSE) measures whether multi-view consistency is preserved over time. It is the Sampson error across different generated views averaged in time. The fundamental matrix used in CSE is estimated using the keypoints accumulated across all temporal frames.
如表 24 所示,我们发现 Cosmos-Predict1-7B-Text2World-Sample-MultiView 和 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 都比 VideoLDM-MultiView 展现出更好的多视角几何一致性。对于从我们的 WFM 微调得到的世界模型,生成视频的整体几何合理性要好得多。我们还注意到,在轨迹控制条件下,由于显式的 3D 引导,这种一致性进一步提高,Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 排名最佳。
As shown in Table 24, we find that both Cosmos-Predict1-7B-Text2World-Sample-MultiView and Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond render better multi-view geometry consistency over VideoLDM-MultiView. The overall geometric plausibility of the generated videos is much better for a world model fine-tuned from our WFM. We also note that, with the trajectory control condition, such consistency is further improved thanks to the explicit 3D guidance as Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond is ranked the best.
轨迹一致性:轨迹一致性误差(TAE)。我们设计了一个鲁棒的多视角相机位姿估计流程,类似于基于公式的流程。该位姿估计流程具有在线动态掩码生成模块和高效密集光束法平差模块,能够实现鲁棒且实时的多视角相机位姿估计。我们使用该流程分别使用两种多视角相机配置(“前 + 左前”相机和“前 + 右前”相机)来估计前置相机的位姿。然后计算它们的轨迹误差以显示它们的一致性,反映多视角生成的一致性。具体来说,我们计算绝对轨迹误差(ATE)和相对位姿误差的平移分量(RPE-t)和旋转分量(RPE-R)。我们将轨迹长度归一化为 1.0 以进行公平比较,并排除相机运动较小的情况(例如,汽车在红灯前停下)。
Trajectory consistency: Trajectory Agreement Error (TAE). We design a robust multi-view camera pose estimation pipeline similar to the one used in based on the formulation of . Such a pose estimation pipeline features an online dynamic mask generation module and a highly efficient dense bundle adjustment module, reaching a robust and real-time performance for estimating multi-view camera poses. We use this pipeline to estimate the camera poses of the front camera, separately using two multi-view camera configurations that consider “front + left-front” cameras and “front + right-front” cameras. We then calculate their trajectory errors to show their agreement, reflecting the consistency of the multi-view generation. Specifically, we compute the Absolute Trajectory Error (ATE) and Relative Pose Error for both the translational component (RPE-t) and the rotational component (RPE-R). We normalize the length of the trajectories to 1.0 for fair comparison and exclude the cases with minor camera movements (e.g., a car stopped at a red light).
如表 25 所示,结果与多视角几何一致性的发现相呼应,从 Cosmos WFM 微调得到的世界模型的轨迹一致性远好于 VideoLDM-MultiView。我们注意到,后训练的 Cosmos 世界模型的轨迹一致性接近真实世界视频。
As shown in Table 25, the results echo the findings from the multi-view geometry consistency, where the trajectory consistency of the world models fine-tuned from the Cosmos WFM is much better than that from the VideoLDM-MultiView. We note that the post-trained Cosmos world models have trajectory consistency that is close to real-world videos.
轨迹一致性:轨迹跟随误差(TFE)。此外,对于将轨迹控制条件输入到模型中的模型,我们使用上述相同的相机位姿估计流程,利用多视角信息计算前置相机的位姿,并将预测轨迹与真实轨迹条件进行比较。这衡量了模型遵循给定轨迹路径的程度。如表 25 所示,使用我们的 Cosmos 后训练世界模型生成的视频估计的轨迹误差仅比真实值基准低 7 厘米以内。如此微小的差距表明我们的模型能够准确遵循给定的轨迹路径,这对于训练自动驾驶智能体至关重要。
Trajectory consistency: Trajectory Following Error (TFE). Furthermore, for the model where we have trajectory control conditions being fed into the model, we use the same camera pose estimation pipeline used above to compute the poses of the front camera using multi-view information and compare the predicted trajectory with respect to the ground truth trajectory condition. This measures how well the model follows the given trajectory path. As shown in Table 25, the trajectory error estimated using the generated videos from our Cosmos post-trained world models is only \(<\) 7cm less precise than the ground-truth oracle. Such a considerably slight margin shows that our model is able to accurately follow the given trajectory path, which is crucial for training autonomous driving agents.
**目标跟踪一致性**。最后,我们使用 YOLOv11x 对生成的 8 秒视频进行了目标检测与跟踪。人类标注员的任务是识别跟踪算法误解物理上不可能场景的情况,例如两个不同的物体(如一个人和一辆车)错误地合并为单个跟踪实体。为了评估这一点,我们向标注员提供了包含 157 个物体的 20 个生成视频的随机样本。值得注意的是,157 个物体中没有一个出现物理上不可能的场景,这证明了我们生成的驾驶视频的物理一致性和物体恒存性。
Objects tracking consistency. Finally, we applied object detection and tracking using YOLOv11x on the generated 8-second videos. Human annotators were tasked with identifying instances where the tracking algorithm misinterpreted physically impossible scenarios, such as two distinct objects (e.g., a person and a car) merging incorrectly into a single tracked entity. To evaluate this, we provided annotators with a random sample of 20 generated videos containing 157 objects. Remarkably, none of the 157 objects exhibited any physically impossible scenarios, demonstrating the physical consistency and object permanence of our generated driving videos.
为了安全使用我们的世界基础模型(WFMs),我们开发了一套全面的护栏系统。该系统包括两个阶段:前置护栏阶段和后置护栏阶段。前置护栏阶段利用 Aegis 和关键词列表来阻止有害提示。后置护栏阶段使用视频内容安全分类器和人脸模糊过滤器来阻止有害的视觉输出。该流程如图 30 所示。
For the safe use of our WFMs, we develop a comprehensive guardrail system. It consists of two stages: the pre-Guard stage and the post-Guard stage. The pre-Guard stage leverages Aegis and a keyword list to block harmful prompts. The post-Guard stage blocks harmful visual outputs using a video content safety classifier and a face blur filter. The pipeline is illustrated in Fig. 30.
我们的前置守卫是一种文本域护栏,由基于 LLM 的护栏(用于处理语义复杂的提示)和基于简单黑名单的检查器(用于明确不安全的关键词)组成。
Our pre-Guard is a text-domain guardrail comprising an LLM-based guardrail for semantically complex prompts and a simple blocklist-based checker for explicitly unsafe keywords.
黑名单启发式方法将作为第一道防线,以降低生成不安全内容的风险。该方法通过对提示词进行关键词搜索,与一个硬编码的黑名单(包含大量露骨和不当词汇的语料库)进行匹配,从而阻止明显有害的生成内容。输入词汇使用 WordNetLemmatizer 进行词形还原,该工具利用英语词汇数据库提取单词的变体形式对应的词根。例如,“abacii”的词根是“abacus”。然后将这些词形还原后的单词与硬编码黑名单中的词汇进行比较,如果发现任何不当词汇,则整个提示词将被拒绝。我们使用全面的关键词集,以最大程度地保护用户。
The blocklist heuristic will act as the first line of defense to mitigate the risk of generating unsafe content. This is designed to block explicitly harmful generations by doing a keyword search on the prompt against a hard-coded blocklist of a large corpus of explicit and objectionable words. Input words are lemmatized using WordNetLemmatizer, a tool that uses a lexical database of the English language to extract the root word from its variants. For example, the root word of “abacii” is “abacus”. These lemmatized words are then compared to the words in the hard-coded blocklist, and the entire prompt is rejected if any profanity is found. We use a comprehensive set of keywords to maximally protect our users.
作为第二道防线,我们使用 Aegis-AI-Content-Safety-LlamaGuard-LLM-Defensive-1.0,这是 Llama-Guard 的一个微调版本,基于 NVIDIA 的 Aegis 内容安全数据集训练,覆盖 NVIDIA 广泛分类法中的 13 个关键安全风险类别。AEGIS 1.0 有两个版本:防御版本和许可版本。防御版本采用比许可版本更严格的权限边界。Cosmos 使用防御版本的 Aegis 来阻止试图生成有害内容的潜在有害用户提示。如果输入提示被此提示过滤器归类为不安全,则不会生成视频,并显示错误消息。
As the second line of defense, we use Aegis-AI-Content-Safety-LlamaGuard-LLM-Defensive-1.0, which is a fine-tuned version of Llama-Guard trained on NVIDIA's Aegis Content Safety Dataset covering NVIDIA's broad taxonomy of 13 critical safety risk categories. There are two versions of AEGIS 1.0: the defensive version and the permissive version. The defensive version adopts a tighter permission boundary than the permissive version. Cosmos uses the defensive version of Aegis to block potentially harmful user prompts that attempt to generate harmful content. If the input prompt is categorized as unsafe by this prompt filter, the video is not generated, and an error message is displayed.
为了将 Aegis 用作提示过滤器,如果提示属于以下类别,我们将其归类为不安全:暴力、性、犯罪计划、武器、药物滥用、自杀、儿童性虐待材料、仇恨、骚扰、威胁和亵渎。从提示过滤的角度来看,任何不属于上述类别的提示均被视为安全。
For using Aegis as a prompt filter, we classify the prompt as unsafe if it falls into the following categories: violence, sexual, criminal planning, weapons, substance abuse, suicide, child sexual abuse material, hatred, harassment, threat, and profanity. Any prompt that does not fall into the above categories is considered safe from the prompt-filtering standpoint.
我们的后置防护是一个视觉域护栏,包含视频内容安全过滤器和针对生成输出的人脸模糊过滤器。
Our post-Guard is a vision-domain guardrail comprising a video content safety filter and a face blur filter for the generated output.
视频内容安全过滤器是一个帧级多类分类器,在我们的视频数据集和生成结果上进行训练。在这些类别中,有些被认为是安全的,而另一些则是不安全的。训练该分类器的一个主要挑战是平衡误报(即安全内容被错误标记为不安全)和漏报(即不安全内容被错误分类为安全)。为了最小化分类错误,我们在训练过程中仔细平衡了数据。
The Video Content Safety Filter is a frame-level multi-class classifier trained on our video dataset and generation results. Among the classes, some are considered safe, while others are unsafe. A major challenge in training the classifier is in balancing false positives, where safe content is mistakenly flagged as unsafe, and false negatives, where unsafe content is wrongly classified as safe. To minimize classification errors, we carefully balanced the data during training.
我们收集了三种类型的真实标注数据。首先,我们从数据集中采样大量视频,提取帧,并使用 VLM 确定其类别。其次,我们使用一组提示词通过我们的 WFM 生成合成视频,以确保覆盖边缘案例和代表性不足的内容类别。最后,人工标注员为我们数据集的一部分提供“金标准”标签,增加了至关重要的验证层,并帮助我们不断优化分类器的准确性。我们为每个视频帧提取 SigLIP 嵌入,并在这些嵌入上训练一个简单的 MLP 分类器。在推理过程中,我们为每一帧生成 SigLIP 嵌入,然后应用分类器。如果任何一帧被分类为不安全,则整个视频被标记为不安全。
We collect three kinds of ground truth annotated data. First, we sample a large set of videos from our dataset, extract frames, and determine its class using a VLM. Next, we generate synthetic videos with our WFMs using a set of prompts to ensure coverage of corner cases and least-represented content categories. Finally, human annotators provide the “gold standard” labels for a portion of our dataset, adding a vital layer of validation and helping us continuously refine the accuracy of our classifier. We extract the SigLIP embedding for each video frame and train a simple MLP classifier on the embeddings. During inference, we generate a SigLIP embedding for every frame and then apply the classifier. The entire video is flagged as unsafe if any frame is classified as unsafe.
我们使用 RetinaFace(一种最先进的人脸检测模型)来识别高置信度的人脸区域。对于任何检测到的大于 \(20\times 20\) 像素的人脸区域,我们应用像素化处理以遮蔽这些区域,同时保留整体场景构图,以满足 Physical AI 应用的需求。
We use RetinaFace, a state-of-the-art face detection model, to identify facial regions with high confidence scores. For any detected face region larger than \(20\times 20\) pixels, we apply pixelation to obscure the regions while preserving the overall scene composition for Physical AI applications.
我们设立了一个专门的红队,利用内部攻击提示数据集中收集的标准和对抗性示例,主动探测系统。这些视频输出由一组经过专门训练、针对我们任务的专家标注员进行标注,按照 7.1.2 节中的分类法,对生成的视频在多个危害类别上进行 1-5 分的评分。这些标注还指定了检测到不安全内容的起始帧和结束帧,从而生成高质量的标注。红队还针对每个护栏组件独立进行定向测试,以识别弱点并改进边缘情况下的性能。截至发表日期,红队已测试并标注了超过\(10{,}000\)个精心设计的提示-视频对,覆盖了广泛的不安全内容。
We employ a dedicated red team to actively probe the system using both standard and adversarial examples that are collected in an internal attack prompt dataset. These video outputs are annotated by a team of expert annotators, who were specially trained for our task, to classify the generated video on a scale of 1-5 on multiple categories of harm related to the taxonomy in section 7.1.2. These annotations also specify the start and end-frames where the unsafe content is detected, thereby generating high-quality annotations. The red team also probed each guardrail component independently with targeted examples to identify weaknesses and improve performance in edge cases. As of the date of publication, the red team has tested and annotated over \(10{,}000\) distinct prompt-video pairs that were carefully crafted to cover a broad range of unsafe content.
世界模型。“世界模型”的概念源于 Ha 和 Schmidhuber 的开创性工作,该工作提出利用神经网络模型学习真实世界的表示,以根据当前状态和输入预测未来状态。物理世界模型的准确表示不仅能够可靠地预测未来状态,还能为决策提供信息。这种对物理世界建模的概念并不新鲜;传统的自动化和机器人行业早已在规划和控制算法中采用基于物理定律和系统辨识的数学模型。然而,这些特定于系统的模型通常局限于低维状态空间,限制了跨系统的泛化和知识迁移,在应用于新任务或新环境时限制了模型的复用。近年来深度学习的进展,特别是生成式 AI,使得直接从视觉观察中学习世界模型成为可能。
World models. The concept of "world models" originated from the seminal work of Ha and Schmidhuber, which proposed learning a representation of the real world using neural network models to predict future states given current states and inputs. An accurate representation of the physical world model enables not only reliable prediction of future states but also informed decision-making. This concept of modeling the physical world is not new; traditional automation and robotics industries have long employed mathematical models based on physics laws and system identification in planning and control algorithms. However, these system-specific models, typically confined to low-dimensional state spaces, restrict generalization and knowledge transfer across different systems, limiting model reuse when applied to new tasks or environments. Recent advances in deep learning, particularly generative AI, have made it possible to learn world models directly from visual observations.
现代世界模型流程可根据其骨干架构进行分类。大多数工作,包括 Ha 和 Schmidhuber 的原始论文,采用循环神经网络对通过自编码器学习的潜在空间中的系统状态演化进行建模。更近的趋势将世界模型视为视觉观察空间中的生成模型,通常采用条件视频生成模型的形式(例如,动作到视频、文本到视频)。这些模型可以是自回归的或基于扩散的,正如本工作所考虑的。另一种有前景的方法是生成式仿真,它结合生成式 AI 和物理模拟器来建模真实世界。
Modern world model pipelines can be categorized based on their backbone architecture. Most works, including the original paper by Ha and Schmidhuber, employ a recurrent neural network to model system state evolution in a latent space learned via an autoencoder. More recent trends view world models as generative models in visual observation space, often in the form of conditional video generative models (e.g., action-to-video, text-to-video). These models can be either autoregressive or diffusion-based, as considered in this work. Another promising approach is generative simulation, which combines generative AI and physical simulators to model the real world.
训练良好的世界模型可以以多种方式应用,包括验证、基于规划的模型预测控制以及基于模型的强化学习。世界模型的有效性已在计算机游戏、真实世界机器人和自动驾驶等领域得到验证。我们设想基础世界模型将对这些行业产生变革性影响。
A well-trained world model can be applied in various ways, including verification, planning-based model predictive control, and model-based reinforcement learning. The effectiveness of world models has been demonstrated in domains such as computer games, real-world robots, and autonomous driving. We envision that foundational world models will have transformative impacts on these industries.
视频生成模型。近年来,视频生成模型领域经历了快速发展。从最初生成短时长、低分辨率视频的模型,该领域已显著发展,视频生成模型现已成为生成式 AI 研究的前沿。近年来涌现出令人瞩目的视频生成模型,如 Sora、Dream Machine、Gen 3 和 Kling,能够生成逼真的高分辨率视频。这些进展是在首个视频生成模型发布后的短短几年内取得的。
Video generative models. The field of video generative models has undergone rapid development in recent years. From the initial models that produced short, low-resolution videos, the field has evolved significantly, with video generative models now at the forefront of generative AI research. Recent years have seen the emergence of impressive video generative models, such as Sora, Dream Machine, Gen 3 and Kling, capable of producing realistic, high-resolution videos. These advancements have been made in just a few years since the release of the first video generative model.
现有的大多数视频生成模型工作集中于文本到视频任务,即根据文本提示输入生成视频。这些模型使用户能够通过精心设计的文本提示创建令人印象深刻的视频。其他流行的任务包括图像到视频(从给定图像帧生成视频)、视频到视频(根据参考视频生成新视频)以及动作到视频(基于动作生成视频),后者受世界模型和具身 AI 发展的驱动。
Most existing work on video generative models focuses on text-to-video tasks, which generate videos based on text prompt inputs. These models enable users to create impressive videos using carefully designed text prompts. Other popular tasks include Image-to-Video that generates videos starting from a given image frame, Video-to-Video that generates new videos given a reference video, and Action-to-Video that generates videos based on actions driven by the development of world models and embodied AI.
大多数视频生成模型采用扩散模型框架,逐步将噪声转化为视频序列。自回归模型也被用于视频生成,其优势在于能够以统一的方式处理视频和其他模态。尽管自回归模型已显示出潜力,但基于扩散的视频模型在视觉质量上仍然更胜一筹。我们的目标是帮助物理 AI 开发者推进其应用。我们认为,基于扩散和基于自回归的模型各有优缺点。基于扩散的模型能生成视觉质量更好的视频,而基于自回归的模型能更好地利用 LLM 社区开发的各种技术。我们构建了基于扩散的(Cosmos-Diffusion)和基于自回归的(Cosmos-Autoregressive)世界基础模型(WFM),并将其提供给物理 AI 构建者。
The majority of video generative models adopt the diffusion model framework to gradually transform noise into video sequences. Autoregressive models have also been employed for video generation, offering the advantage of handling video and other modalities in a unified manner. While autoregressive models have shown promise, diffusion-based video models still excel in terms of visual quality. Our goal is to help Physical AI developers advance their applications. We believe that the diffusion-based and autoregressive-based models both have their pros and cons. Diffusion-based models can render videos with better visual quality. Autoregressive-based models can better leverage all sorts of techniques developed by the LLM community. We build both diffusion-based (Cosmos-Diffusion) and autoregressive-based (Cosmos-Autoregressive) WFMs and make them available to the Physical AI builders.
带相机控制的视频生成。3D 一致的视频生成可追溯到视图合成和 3D 重建的早期工作,当时社区试图利用神经渲染应用于各种 3D 表示来创建 3D 一致的视频。在这一研究路线中,单图像 3D 视图合成尤其具有挑战性,通常需要从多视图图像数据集中学习强大的 3D 先验模型。由于此类 3D 先验模型往往难以良好扩展,基于学习的视图合成也通过使用可扩展的 Transformer 架构的纯数据驱动方法进行了探索。这绕过了对显式 3D 先验知识的需求:不再依赖应用于 3D 表示的神经渲染,而是由以相机输入为条件的神经网络直接合成新视图。这一范式已通过扩散模型成功扩展,并在 3D 资产生成中得到广泛应用。近期视频生成质量的进展表明,通过扩展训练视频数据规模,有可能实现完全的 3D 一致性。此类模型上的相机可控性因其在机器人和自主导航中的巨大应用潜力,已成为一个活跃的研究领域。
Video generation with camera control. 3D-consistent video generation traces back to early works in view synthesis and 3D reconstruction, where the community sought to create 3D-consistent videos using neural rendering applied to various 3D representations. Within this line of research, single-image 3D view synthesis is particularly challenging, typically requiring the learning of a strong 3D prior model from multi-view image datasets. As such 3D prior models often suffer to scale well, learning-based view synthesis has also been explored through a purely data-driven approach using scalable Transformer architectures. This bypasses the need for explicit 3D prior knowledge: instead of relying on neural rendering applied to 3D representations, novel views are synthesized directly by neural networks conditioned on camera inputs. This paradigm has been successfully scaled up with diffusion models, finding broad applications in 3D asset generation. Recent advances in video generation quality suggest the potential for achieving full 3D consistency through the scaling of training video data. Camera controllability on such models has since become an active area of investigation for its great potential applications to robotics and autonomous navigation.
用于机器人控制的生成模型。深度生成模型的最新进展引发了人们对其在机器人控制中应用的极大兴趣。已经出现了多种方法,其中一条工作路线直接将扩散模型用作视觉运动策略,在多种机器人任务中展示了模仿学习的显著改进。另外两条与本工作更相关的路线是:使用预训练的图像和视频生成模型作为运动规划器,以及使用图像和视频数据进行生成式预训练。生成式运动规划方法旨在通过生成中间视觉子目标而非显式动作序列来增强对未见环境的泛化能力。这种视觉表示策略被证明更为鲁棒,因为图像和视频子目标可以跨不同环境设置进行泛化,而动作序列通常特定于环境和任务。生成式预训练方法利用大规模图像和视频数据集进行预训练。其中,从预训练的文本到图像扩散模型中提取并利用特征来指导后续的策略学习,并使用两阶段框架:首先预训练模型以预测未来帧,然后微调以联合预测动作和未来帧。
Generative models for robotic control. Recent advances in deep generative models have sparked significant interest in their application to robotic control. Several approaches have emerged, with one line of work directly employing diffusion models as visuomotor policies, demonstrating substantial improvements in imitation learning in various robotic tasks. Two other threads more related to this work are the use of pre-trained image and video generation models as motion planners and the use of image and video data for generative pre-training. The generative motion planning approach aims to enhance generalization to unseen environments by generating intermediate visual sub-goals rather than explicit action sequences. This visual representation strategy proves more robust, as image and video sub-goals can generalize across diverse environmental setups, unlike action sequences that are typically environment- and task-specific. The generative pre-training approach leverages large-scale image and video datasets for pre-training. While extract and utilize features from pre-trained text-to-image diffusion models to guide subsequent policy learning, and use a two-stage framework: first pre-training the model to predict future frames, then fine-tuning it to jointly predict both actions and future frames.
用于自动驾驶的生成模型。视频生成模型有潜力通过生成以多种输入模态(如文本、图像、轨迹、3D 数据或地图)为条件的逼真驾驶视频,彻底改变自动驾驶模拟。尽管潜力巨大,现有方法受限于数据规模、分辨率和相机视图数量的约束,限制了其作为全面驾驶世界模拟器的有效性。为克服这些限制,我们利用强大的预训练 WFM 的能力,开发了一个灵活且可扩展的驾驶模拟器。我们的模型实现了高分辨率、高帧率和多视图一致性。
Generative models for autonomous driving. Video generative models have the potential to revolutionize autonomous driving simulation by enabling the generation of realistic driving videos conditioned on diverse input modalities, such as text, images, trajectories, 3D data, or maps. Despite their potential, existing approaches have been limited by constraints in data scale, resolution, and the number of camera views, restricting their effectiveness as comprehensive driving world simulators. To overcome these limitations, we leverage the capabilities of a powerful pre-trained WFM to develop a flexible and scalable driving simulator. Our model achieves high resolution, elevated frame rates, and multi-view consistency.
分词器。学习再现输入视觉数据的潜在特征已有相当长的历史。近年来,这类模型(也称为分词器)已被广泛用作提高大规模生成模型训练效率的基本组件。
Tokenizer. There has been a fairly long history of learning latent features that reproduce the input visual data. Recently, such models, also known as tokenizers, have been widely incorporated as essential components to improve the efficiency of training large-scale generative models.
连续视觉分词器,通常包括自编码器(AE)和变分自编码器(VAE),将视觉数据压缩到连续潜空间中,使得基于扩散的模型能够高效训练。在推理时,生成潜变量通过分词器解码器解码回 RGB 空间。多种扩散模型已以这种方式训练用于图像和视频生成。
Continuous visual tokenizers, often including Autoencoder (AE) and Variational Autoencoder (VAE), compress visual data into a continuous latent space where diffusion-based models can be efficiently trained on. At inference time, the generated latents are decoded back to RGB space with the tokenizer decoder. Various diffusion models have been trained in such a way for image and video generation.
离散视觉分词器额外包含一个量化器,将连续潜变量进一步离散化为离散空间,从而便于与文本、音频等其他模态一起集成到大语言模型(LLM)和视觉语言模型(VLM)中。因此,离散分词器被部署在各种视觉理解以及图像和视频生成任务中。
Discrete visual tokenizers additionally involve a quantizer that further discretizes the continuous latents into a discrete space, allowing for easy integration into large language models (LLMs) and vision language models (VLMs) alongside other modalities, such as text and audio. Thus, discrete tokenizers are deployed in various visual understanding as well as image and video generation tasks.
Cosmos 分词器广泛基于先前研究构建,例如 FSQ 和因果架构,旨在创建一套高效且高质量的分词器。
Cosmos tokenizers are extensively built based on previous studies, e.g., FSQ and causal architecture, with the goal of creating a suite of efficient and high-quality tokenizers.
Cosmos 世界基础模型标志着向构建物理世界通用模拟器迈出的重要一步。本文概述了我们的综合方法,包括数据整理流程、连续和离散分词器的设计、扩散和自回归世界基础模型的架构,以及针对各种下游物理 AI 任务的微调过程。值得注意的是,我们展示了预训练世界模型对关键应用的适应性,包括 3D 世界导航、机器人操作和自动驾驶系统,这些应用要求 3D 一致性和动作可控性。
The Cosmos World Foundation Models mark a significant step towards building general-purpose simulators for the physical world. This work outlines our comprehensive approach, including the data curation pipeline, the design of continuous and discrete tokenizers, the architecture of diffusion and autoregressive world foundation models, and the fine-tuning process for diverse downstream Physical AI tasks. Notably, we demonstrate the adaptability of our pre-trained world models to critical applications, including 3D world navigation, robotic manipulation, and autonomous vehicle systems, which demand both 3D consistency and action controllability.
局限性。尽管取得了进展,世界基础模型的发展仍处于早期阶段。当前的模型(包括我们的模型)作为物理世界的可靠模拟器仍有所欠缺。我们观察到,我们的模型仍然存在一些问题,包括缺乏物体恒存性、接触丰富动力学的不准确性以及指令遵循的不一致性。此外,生成视频的真实感并不总是反映对基本物理原理(如重力、光相互作用和流体动力学)的遵循。
Limitations. Despite the progress, the development of world foundation models is still in the early stages. Current models, including ours, fall short as reliable simulators of the physical world. We observe that our models still suffer from issues, including the lack of object permanence, inaccuracies in contact-rich dynamics, and inconsistency in instruction following. Additionally, the realism of the generated videos does not always reflect adherence to fundamental physical principles, such as gravity, light interactions, and fluid dynamics.
评估是另一个重大挑战。为人类定义稳健的评估标准来评估物理保真度是困难的,因为此类评估往往受到个人偏见、背景和其他主观因素的影响。此外,这些评估可能不与下游物理 AI 任务中使用的指标正相关。为了解决这些挑战,有前景的方向包括开发由多模态 LLM 驱动的自动评估器,以及利用现有的物理模拟器实现可复现和交互式评估,从而减少对人类评估的依赖。
Evaluation presents another significant challenge. Defining robust rubrics for humans to evaluate physical fidelity is hard as such assessments are often influenced by personal biases, backgrounds, and other subjective factors. Moreover, these evaluations may not align positively with metrics used in downstream Physical AI tasks. In order to address these challenges, promising directions include the development of automated evaluators powered by multi-modal LLMs and leveraging existing physical simulators to enable reproducible and interactive evaluation, thereby reducing dependence on human evaluation.
自回归与扩散 WFM。我们在 3D 一致性(第 5.3.1 节)和机器人视频生成(第 6.2 节)方面的评估结果表明,基于扩散的 WFM 目前能提供更好的生成质量。通过微调,基于扩散的 WFM 能够整合多种控制信号,包括相机姿态、末端执行器位置或自动驾驶车辆轨迹,并生成多视角视频等新颖格式的输出。然而,基于自回归的 WFM 具有巨大的未开发潜力。它们可以(1)利用大型语言模型(LLM)的预训练权重来继承广泛的世界知识,(2)通过使用为因果注意力设计的先进推理优化技术实现更快的生成。如果这些能力得到充分实现,自回归 WFM 可能特别适合需要交互控制或实时处理的应用,例如机器人中的规划和模拟。重要的是,扩散模型和自回归模型之间的界限并非固定不变。最近的进展表明,具有双向注意力的扩散 Transformer 可以蒸馏为具有因果注意力的学生 Transformer,从而在推理过程中支持键值缓存。类似地,自回归模型可以结合局部双向注意力,通过扩散头生成图像。探索这些混合方法及其权衡仍是一个活跃且有前景的研究领域。我们计划进一步研究这些公式,并在未来的工作中提供全面的分析。
Autoregressive vs. Diffusion WFMs. Our evaluation results in 3D consistency (Section 5.3.1) and video generation for robotics (Section 6.2) indicate that diffusion-based WFMs currently deliver better generation quality. Through fine-tuning, diffusion-based WFMs are able to incorporate diverse control signals, including camera pose, end-effector positions, or autonomous vehicle trajectories, and generate outputs of novel formats like multi-view videos. However, autoregressive-based WFMs possess significant untapped potential. They could (1) leverage pre-trained weights from large language models (LLMs) to inherit extensive world knowledge and (2) enable faster generation through the use of advanced inference optimization techniques designed for causal attention. If these capabilities are fully realized, autoregressive WFMs may become particularly well-suited for applications requiring interactive control or real-time processing, such as planning and simulation in robotics. Importantly, the boundary between diffusion and autoregressive models is not rigid. Recent advancements have shown that diffusion transformers with bidirectional attention can be distilled into student transformers with causal attention, enabling support for key-value caching during inference. Similarly, autoregressive models can incorporate locally bidirectional attention to generate images via diffusion heads. Exploring these hybrid approaches and their trade-offs remains an active and promising area of research. We plan to investigate these formulations further and provide a comprehensive analysis in future work.
数据整理 Jacob Huffman、Francesco Ferroni、Alice Luo、Niket Agarwal、Hao Wang、Jing Zhang、David Page、Vasanth Rao Naik Sabavat、Sriharsha Niverty、Erik Barker、Lindsey Pavao、Stella Shi、Prithvijit Chattopadhyay、Shitao Tang、Yin Cui、Yunhao Ge、Qianli Ma、Yifan Ding、Seungjun Nah、Siddharth Gururani、Jiashu Xu、Grace Lam、Tiffany Cai、Jibin Varghese、Pooya Jannaty、Jay Zhangjie Wu、Yuxuan Zhang、Huan Ling、Hanzi Mao、Heng Wang
Data Curation Jacob Huffman, Francesco Ferroni, Alice Luo, Niket Agarwal, Hao Wang, Jing Zhang, David Page, Vasanth Rao Naik Sabavat, Sriharsha Niverty, Erik Barker, Lindsey Pavao, Stella Shi, Prithvijit Chattopadhyay, Shitao Tang, Yin Cui, Yunhao Ge, Qianli Ma, Yifan Ding, Seungjun Nah, Siddharth Gururani, Jiashu Xu, Grace Lam, Tiffany Cai, Jibin Varghese, Pooya Jannaty, Jay Zhangjie Wu, Yuxuan Zhang, Huan Ling, Hanzi Mao, Heng Wang
分词器 Jinwei Gu、Xian Liu、Songwei Ge、Ting-Chun Wang、Haoxiang Wang、Fitsum Reda
Tokenizer Jinwei Gu, Xian Liu, Songwei Ge, Ting-Chun Wang, Haoxiang Wang, Fitsum Reda
基于扩散的世界基础模型预训练 Qinsheng Zhang、Lin Yen-Chen、Xiaohui Zeng、Huan Ling、Shitao Tang、Maciej Bala、Ting-Chun Wang、Yu Zeng、Seungjun Nah、Qianli Ma、Hanzi Mao
Diffusion-based World Foundation Model Pre-training Qinsheng Zhang, Lin Yen-Chen, Xiaohui Zeng, Huan Ling, Shitao Tang, Maciej Bala, Ting-Chun Wang, Yu Zeng, Seungjun Nah, Qianli Ma, Hanzi Mao
基于自回归的世界基础模型预训练 Haoxiang Wang、Yifan Ding、Xian Liu、Jiaojiao Fan、Xiaohui Zeng、Yogesh Balaji
Autoregressive-based World Foundation Model Pre-training Haoxiang Wang, Yifan Ding, Xian Liu, Jiaojiao Fan, Xiaohui Zeng, Yogesh Balaji
提示上采样器 Yunhao Ge、Haoxiang Wang、Jiashu Xu、Yin Cui
Prompt Upsampler Yunhao Ge, Haoxiang Wang, Jiashu Xu, Yin Cui
扩散解码器:Huan Ling、Jiaojiao Fan、Fitsum Reda、Yogesh Balaji、Hanzi Mao、Qinsheng Zhang
Diffusion Decoder: Huan Ling, Jiaojiao Fan, Fitsum Reda, Yogesh Balaji, Hanzi Mao, Qinsheng Zhang
3D 一致性预训练评估:Jiahui Huang、Chen-Hsuan Lin
3D Consistency Pre-training Evaluation: Jiahui Huang, Chen-Hsuan Lin
物理对齐预训练评估:Francesco Ferroni、Prithvijit Chattopadhyay、Xinyue Wei、Qianli Ma、Gergely Klár、Chen-Hsuan Lin
Physics Alignment Pre-training Evaluation: Francesco Ferroni, Prithvijit Chattopadhyay, Xinyue Wei, Qianli Ma, Gergely Klár, Chen-Hsuan Lin
相机控制后训练评估:Xiaohui Zeng、Tsung-Yi Lin、Jingyi Jin、Chen-Hsuan Lin
Camera Control Post-training Evaluation: Xiaohui Zeng, Tsung-Yi Lin, Jingyi Jin, Chen-Hsuan Lin
机器人后训练评估:Lin Yen-Chen、Wei-Cheng Tseng、Yunhao Ge、Xian Liu、Shitao Tang、Fangyin Wei、Lyne Tchapmi、Yu Zeng、Qingqing Zhao、Yin Cui、Zhaoshuo Li、Jinwei Gu
Robotics Post-training Evaluation: Lin Yen-Chen, Wei-Cheng Tseng, Yunhao Ge, Xian Liu, Shitao Tang, Fangyin Wei, Lyne Tchapmi, Yu Zeng, Qingqing Zhao, Yin Cui, Zhaoshuo Li, Jinwei Gu
自动驾驶后训练评估:Seung Wook Kim、Jay Zhangjie Wu、Jiahui Huang、Francesco Ferroni、Michele Fenzi、Daniel Dworakowski、Despoina Paschalidou、Ed Schmerling、Shiyi Lan、Laura Leal-Taixe、Sanja Fidler、Huan Ling
Autonomous Driving Post-training Evaluation: Seung Wook Kim, Jay Zhangjie Wu, Jiahui Huang, Francesco Ferroni, Michele Fenzi, Daniel Dworakowski, Despoina Paschalidou, Ed Schmerling, Shiyi Lan, Laura Leal-Taixe, Sanja Fidler, Huan Ling
护栏:Jibin Varghese、Arslan Ali、Grace Lam、Pooya Jannaty
Guardrail: Jibin Varghese, Arslan Ali, Grace Lam, Pooya Jannaty
平台架构师:Ming-Yu Liu
Platform Architect: Ming-Yu Liu
Anqi Li, Arsalan Mousavian, Artur Zolkowski, Bartosz Stefaniak, Dieter Fox, Ethan He, Kaichun Mo, Morteza Ramezanali, Przemek Tredak, Wei Yang, Xiaowei Ren, Yongxin Chen, Zeeshan Patel
Anqi Li, Arsalan Mousavian, Artur Zolkowski, Bartosz Stefaniak, Dieter Fox, Ethan He, Kaichun Mo, Morteza Ramezanali, Przemek Tredak, Wei Yang, Xiaowei Ren, Yongxin Chen, Zeeshan Patel
我们感谢 1X Technologies 慷慨提供人形机器人数据,并为本技术报告中机器人操作的后训练提供了宝贵的支持。
We thank 1X Technologies for generously providing humanoid robot data and offering invaluable support for the post-training for robotic manipulation in this technical report.
我们感谢 Aarti Basant、Akan Huang、Alex Qi、Alexis Bjorlin、Amanda Moran、Amol Fasale、Ankit Patel、Arash Vahdat、Aryaman Gupta、Ashna Khetan、Ashwath Aithal、Bor-Yiing Su、Bryan Catanzaro、Charles Hsu、Chris Pruett、Christopher Horvath、Clark Doan、Coulten Holt、Dane Aconfora、Deepak Narayanan、Dennis Chang、Dheeraj Kapur、Dong Ahn、Ebrar Erdem、Elmar Haussmann、Fuzhao Xue、Gandhi Vaithilingam、Henry Estela、Henry Vera、Herb Woodruff、Imad El Hanafi、Jashojit Mukherjee、Jason Sewall、Jensen Huang、John Dickinson、Jonah Alben、Jonah Philion、Josh Abbott、Jun Gao、Kumar Anik、Lee Ditiangkin、Ligeng Zhu、Linxi Fan、Luke Alonso、Madison Huang、Marek Dabek、Mark Arnold、Max Ehrlich、Michele Ferretti、Misbah Mubarak、Misha Smelyanskiy、Mohamed Fawzy、Mohammad Harrim、Mohammad Shoeybi、Omkar Mehta、Pallab Bhattacharya、Paniz Karbasi、Pasha Shamis、Raju Wagwani、Rick Izzo、Robert Hero、Sharon Clay、Song Han、Songyan Tang、Sophia Huang、Sridhar Bhuvanapalli、TJ Galda、Thomas Volk、Tobias Lasser、Vaibhav Ranglani、Vijay Anand Korthikanti、Yao Lu、Yazdan Aghaghiri、Yugi Guvvala、Yuke Zhu 和 Zekun Hao 的反馈和工程支持。
We thank Aarti Basant, Akan Huang, Alex Qi, Alexis Bjorlin, Amanda Moran, Amol Fasale, Ankit Patel, Arash Vahdat, Aryaman Gupta, Ashna Khetan, Ashwath Aithal, Bor-Yiing Su, Bryan Catanzaro, Charles Hsu, Chris Pruett, Christopher Horvath, Clark Doan, Coulten Holt, Dane Aconfora, Deepak Narayanan, Dennis Chang, Dheeraj Kapur, Dong Ahn, Ebrar Erdem, Elmar Haussmann, Fuzhao Xue, Gandhi Vaithilingam, Henry Estela, Henry Vera, Herb Woodruff, Imad El Hanafi, Jashojit Mukherjee, Jason Sewall, Jensen Huang, John Dickinson, Jonah Alben, Jonah Philion, Josh Abbott, Jun Gao, Kumar Anik, Lee Ditiangkin, Ligeng Zhu, Linxi Fan, Luke Alonso, Madison Huang, Marek Dabek, Mark Arnold, Max Ehrlich, Michele Ferretti, Misbah Mubarak, Misha Smelyanskiy, Mohamed Fawzy, Mohammad Harrim, Mohammad Shoeybi, Omkar Mehta, Pallab Bhattacharya, Paniz Karbasi, Pasha Shamis, Raju Wagwani, Rick Izzo, Robert Hero, Sharon Clay, Song Han, Songyan Tang, Sophia Huang, Sridhar Bhuvanapalli, TJ Galda, Thomas Volk, Tobias Lasser, Vaibhav Ranglani, Vijay Anand Korthikanti, Yao Lu, Yazdan Aghaghiri, Yugi Guvvala, Yuke Zhu and Zekun Hao for their feedback and engineering support.
我们感谢 Iain Cunningham、Jim Fan、Marco Pavone、Meredith Price、Nikki Pope 和 Scott Reed 对本技术报告早期草稿的反馈。
We thank Iain Cunningham, Jim Fan, Marco Pavone, Meredith Price, Nikki Pope and Scott Reed for their feedback on the early draft of this technical report.