V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→现代人工智能面临的一个主要挑战是如何主要通过观察来学习理解世界并学会行动。本文探索了一种自监督方法,将互联网规模的视频数据与少量交互数据(机器人轨迹)相结合,以开发能够理解、预测和规划物理世界的模型。我们首先在包含超过 100 万小时互联网视频的视频和图像数据集上,预训练了一个无动作的联合嵌入预测架构 V-JEPA 2。V-JEPA 2 在运动理解(Something-Something v2 上 top-1 准确率 77.3)和人类动作预测(Epic-Kitchens-100 上 recall-at-5 为 39.7)方面取得了强劲性能,超越了之前的任务特定模型。此外,在将 V-JEPA 2 与大型语言模型对齐后,我们在 8 亿参数规模的多项视频问答任务上展示了最先进的性能(例如,PerceptionTest 上 84.0,TempCompass 上 76.9)。最后,我们展示了如何通过后训练一个潜在动作条件世界模型 V-JEPA 2-AC,使用来自 Droid 数据集的不到 62 小时的无标签机器人视频,将自监督学习应用于机器人规划任务。我们在两个不同实验室的 Franka 机械臂上零样本部署了 V-JEPA 2-AC,并实现了使用图像目标进行规划来拾取和放置物体。值得注意的是,这是在未从这些环境中的机器人收集任何数据,且没有任务特定训练或奖励的情况下实现的。这项工作展示了如何通过从网络规模数据和少量机器人交互数据中进行自监督学习,得到一个能够在物理世界中进行规划的世界模型。
A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data (robot trajectories), to develop models capable of understanding, predicting, and planning in the physical world. We first pre-train an action-free joint-embedding-predictive architecture, V-JEPA 2, on a video and image dataset comprising over 1 million hours of internet video. V-JEPA 2 achieves strong performance on motion understanding (77.3 top-1 accuracy on Something-Something v2) and state-of-the-art performance on human action anticipation (39.7 recall-at-5 on Epic-Kitchens-100) surpassing previous task-specific models. Additionally, after aligning V-JEPA 2 with a large language model, we demonstrate state-of-the-art performance on multiple video question-answering tasks at the 8 billion parameter scale (e.g., 84.0 on PerceptionTest, 76.9 on TempCompass). Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the Droid dataset. We deploy V-JEPA 2-AC zero-shot on Franka arms in two different labs and enable picking and placing of objects using planning with image goals. Notably, this is achieved without collecting any data from the robots in these environments, and without any task-specific training or reward. This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.
现代 AI 的一个主要挑战是学会通过观察来理解世界并采取行动。本文探索了一种自监督方法,将互联网规模的视频数据与少量交互数据(机器人轨迹)相结合,以开发能够理解、预测和规划物理世界的模型。我们首先在一个包含超过 100 万小时互联网视频的视频和图像数据集上,预训练了一个无动作的联合嵌入预测架构 V-JEPA 2。V-JEPA 2 在运动理解(Something-Something v2 上 77.3 的 top-1 准确率)和人类动作预测(Epic-Kitchens-100 上 39.7 的 recall-at-5)方面取得了强劲性能,超越了之前的任务特定模型。此外,在将 V-JEPA 2 与大型语言模型对齐后,我们在 8 亿参数规模的多项视频问答任务上展示了最先进的性能(例如,PerceptionTest 上 84.0,TempCompass 上 76.9)。最后,我们展示了如何通过后训练一个潜在动作条件世界模型 V-JEPA 2-AC,使用来自 Droid 数据集的不到 62 小时的无标签机器人视频,将自监督学习应用于机器人规划任务。我们在两个不同实验室的 Franka 机械臂上零样本部署了 V-JEPA 2-AC,并实现了基于图像目标的规划来拾取和放置物体。值得注意的是,这一实现无需从这些环境中的机器人收集任何数据,也无需任何任务特定训练或奖励。这项工作证明了如何从网络规模数据和少量机器人交互数据中进行自监督学习,能够产生一个能够在物理世界中进行规划的世界模型。
A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data (robot trajectories), to develop models capable of understanding, predicting, and planning in the physical world. We first pre-train an action-free joint-embedding-predictive architecture, V-JEPA 2, on a video and image dataset comprising over 1 million hours of internet video. V-JEPA 2 achieves strong performance on motion understanding (77.3 top-1 accuracy on Something-Something v2) and state-of-the-art performance on human action anticipation (39.7 recall-at-5 on Epic-Kitchens-100) surpassing previous task-specific models. Additionally, after aligning V-JEPA 2 with a large language model, we demonstrate state-of-the-art performance on multiple video question-answering tasks at the 8 billion parameter scale (e.g., 84.0 on PerceptionTest, 76.9 on TempCompass). Finally, we show how self-supervised learning can be applied to robotic planning tasks by post-training a latent action-conditioned world model, V-JEPA 2-AC, using less than 62 hours of unlabeled robot videos from the Droid dataset. We deploy V-JEPA 2-AC zero-shot on Franka arms in two different labs and enable picking and placing of objects using planning with image goals. Notably, this is achieved without collecting any data from the robots in these environments, and without any task-specific training or reward. This work demonstrates how self-supervised learning from web-scale data and a small amount of robot interaction data can yield a world model capable of planning in the physical world.
人类在承担新任务和在陌生环境中操作时具有适应和泛化的能力。一些认知学习理论认为,人类通过整合低层感官输入来学习一个内部世界模型,以表示和预测未来状态;这些理论进一步指出,这个世界模型在任何给定时刻塑造我们的感知,在告知我们对现实的理解方面起着关键作用。此外,我们预测自身行为对未来世界状态影响的能力对于目标导向的规划也至关重要。构建从感官数据(如视频)中学习世界模型的人工智能体,可以使它们理解物理世界、预测未来状态,并像人类一样在新情境中有效规划,从而产生能够处理前所未见任务的系统。
Humans have the ability to adapt and generalize when taking on new tasks and operating in unfamiliar environments. Several cognitive learning theories suggest that humans learn an internal model of the world by integrating low-level sensory inputs to represent and predict future states, and they further posit that this world model shapes our perception at any given moment, playing a crucial role in informing our understanding of reality. Moreover, our ability to predict the effects of our actions on future states of the world is also essential for goal-oriented planning. Building artificial agents that learn a world model from sensory data, such as video, could enable them to understand the physical world, predict future states, and effectively — like humans — plan in new situations, resulting in systems capable of tackling tasks that have not been encountered before.
以往的工作探索了从由状态-动作序列组成的交互数据中开发预测性世界模型,通常还依赖来自环境的显式奖励反馈来推断目标。然而,真实世界交互数据的有限可用性限制了这些方法的可扩展性。为了解决这一限制,最近的工作利用互联网规模的视频和交互数据来训练用于机器人控制的动作条件视频生成模型,但在基于模型控制的机器人执行中仅展示了有限的结果。特别是,这类研究往往强调预测保真度和视觉质量的评估,而非规划能力,这可能是由于通过生成视频进行规划的计算成本过高。
Previous works have explored the development of predictive world models from interaction data consisting of state-action sequences, often also relying on explicit reward feedback from the environment to infer goals. However, the limited availability of real-world interaction data constrains the scalability of these methods. To address this limitation, more recent works have leveraged both internet-scale video and interaction data towards training action-conditioned video generation models for robot control, but only demonstrate limited results in robot execution using model-based control. In particular, this line of research often emphasizes the evaluation of the faithfulness of the predictions and visual quality instead of planning capabilities, perhaps due to the computational cost of planning by generating video.
在这项工作中,我们基于自监督假设来学习世界模型,该模型主要从观察中捕获世界的背景知识。具体来说,我们利用联合嵌入预测架构(JEPA),它通过在学习的表示空间中进行预测来学习。与完全从交互数据中学习的方法不同,自监督学习使我们能够利用互联网规模的视频——这些视频描绘了状态序列而没有直接观察动作——来学习表示视频观察,并在该学习表示空间中学习世界动态的预测模型。此外,与基于视频生成的方法不同,JEPA 方法专注于学习场景中可预测方面的表示(例如,运动物体的轨迹),而忽略生成目标所强调的不可预测细节,因为它们进行像素级预测(例如,田野中每片草叶或树上每片叶子的精确位置)。通过扩展 JEPA 预训练,我们证明它产生了具有最先进理解和预测能力的视频表示,并且这些表示可以作为动作条件预测模型的基础,实现零样本规划。
In this work, we build upon the self-supervised hypothesis as a means to learn world models that capture background knowledge of the world largely from observation. Specifically, we leverage the joint-embedding predictive architecture (JEPA), which learns by making predictions in a learned representation space. In contrast to approaches that focus on learning entirely from interaction data, self-supervised learning enables us to make use of internet-scale video — depicting sequences of states without direct observations of the actions — to learn to both represent video observations and learn a predictive model for world dynamics in this learned representation space. Furthermore, in contrast to approaches based on video generation, the JEPA approach focuses on learning representations for predictable aspects of a scene (e.g., the trajectory of an object in motion) while ignoring unpredictable details that generative objectives emphasize, since they make pixel-level predictions (e.g., the precise location of each blade of grass in a field, or each leaf on a tree). By scaling JEPA pretraining, we demonstrate that it yields video representations with state-of-the-art understanding and prediction capabilities, and that such representations can be leveraged as a basis for action-conditioned predictive models and enable zero-shot planning.
我们的方法 V-JEPA 2 采用分阶段训练流程,首先在互联网规模的视频上进行无动作预训练,然后使用少量交互数据进行后训练(见图 1)。在第一阶段,我们使用掩码去噪特征预测目标,其中模型在学习表示空间中预测视频的掩码片段。我们训练 V-JEPA 2 编码器,参数高达 10 亿,视频时长超过 100 万小时。我们的实验证实,扩展自监督视频预训练增强了编码器实现视觉理解的能力,包括广泛的运动和外观识别能力,这通过基于探针的评估以及将编码器与语言模型对齐进行视频问答来验证。
Our approach, V-JEPA 2, utilizes a stage-wise training procedure, beginning with action-free pre-training on internet-scale video, followed by post-training with a small amount of interaction data (see Figure 1). In the first stage, we use a mask-denoising feature prediction objective, where the model predicts masked segments of a video in a learned representation space. We train the V-JEPA 2 encoder with up to 1 billion parameters and with more than 1 million hours of video. Our experiments confirm that scaling self-supervised video pretraining enhances the encoder’s ability to achieve visual understanding, including broad motion and appearance recognition capabilities, through probe-based evaluations and by aligning the encoder with a language model for video question-answering.
在互联网规模视频预训练之后,我们使用第一阶段学习到的表示,在少量交互数据上训练动作条件世界模型 V-JEPA 2-AC。我们的动作条件世界模型是一个 300M 参数的 Transformer 网络,采用块因果注意力机制,它自回归地预测下一视频帧的表示,条件为动作和先前状态。仅使用 Droid 数据集中 62 小时的无标注交互数据,我们证明了训练潜在世界模型的可行性,该模型在给定子目标的情况下,可以用于在 Franka 机器人手臂上规划动作,并在新环境中从单目 RGB 相机零样本执行抓取操作任务。
Following pretraining on internet-scale video, we train an action-conditioned world model, V-JEPA 2-AC, on a small set of interaction data using the representations learned in the first stage. Our action-conditioned world model is a 300M-parameter transformer network employing a block-causal attention mechanism, which autoregressively predicts the representation of the next video frame conditioned on an action and previous states. With as little as 62 hours of unlabeled interaction data from the Droid dataset, we demonstrate the feasibility of training a latent world model that, given sub-goals, can be leveraged to plan actions on a Franka robot arm and perform prehensile manipulation tasks from a monocular RGB camera zero-shot in a new environment.
总而言之,我们展示了基于视频学习的联合嵌入预测架构可用于构建世界模型,从而理解物理世界、预测未来状态并在新情境中有效规划;这是通过利用互联网规模的视频和少量交互数据实现的。具体来说:
To summarize, we show that joint-embedding predictive architectures learning from videos can be used to build a world model that enables understanding the physical world, predicting future states, and effectively planning in new situations; this is achieved by leveraging internet-scale video and a small amount of interaction data. Specifically:
理解——基于探针的分类:扩展自监督视频预训练规模,可得到适用于多种任务的视频表示。V-JEPA 2 擅长编码细粒度的运动信息,在需要运动理解的任务(如 Something-Something v2)上取得了强劲性能,使用注意力探针达到 77.3 的 top-1 准确率。
Understanding — Probe-based Classification: Scaling self-supervised video pretraining results in video representations applicable to many tasks. V-JEPA 2 excels at encoding fine-grained motion information, achieving strong performance on tasks requiring motion understanding, such as Something-Something v2, with 77.3 top-1 accuracy using an attentive probe.
理解——视频问答:V-JEPA 2 编码器可用于训练多模态大语言模型,以解决视频问答任务。我们在多个需要物理世界理解和时间推理的基准上,在 8B 语言模型类别中取得了最先进性能,例如 MVP(44.5 配对准确率)、PerceptionTest(84.0 测试集准确率)、TempCompass(76.9 多选准确率)、TemporalBench(36.7 多二元短问答准确率)和 TOMATO(40.3 准确率)。特别地,我们展示了在没有语言监督的情况下预训练的视频编码器可以与语言模型对齐并取得最先进性能,这与传统观念相反。
Understanding — Video Question-Answering: V-JEPA 2 encoder can be used to train a multi-modal large language model, to tackle video-question answering tasks. We observe state-of-the-art performance on 8B language model class on multiple benchmarks that require physical world understanding and temporal reasoning, such as MVP (44.5 paired accuracy), PerceptionTest (84.0 test set accuracy), TempCompass (76.9 multi-choice accuracy), TemporalBench (36.7 multi-binary short-QA accuracy) and TOMATO (40.3 accuracy). In particular, we show that a video encoder pre-trained without language supervision can be aligned with a language model and achieve state-of-the-art performance, contrary to conventional wisdom.
预测:大规模自监督视频预训练增强了预测能力。V-JEPA 2 在使用注意力探针的 Epic-Kitchens-100 人类动作预期任务上取得了最先进性能,recall-at-5 达到 39.7,相比之前的最佳模型相对提升了 44%。
Prediction: Large-scale self-supervised video pretraining enhances prediction capabilities. V-JEPA 2 achieves state-of-the-art performance on the Epic-Kitchens-100 human-action anticipation task using an attentive probe, with 39.7 recall-at-5, which is a 44% relative improvement over the previous best model.
规划:我们证明了 V-JEPA 2-AC(通过对 V-JEPA 2 进行后训练,仅使用来自流行 Droid 数据集的 62 小时无标注机器人操作数据获得)可以部署在新环境中,通过给定子目标进行规划来解决抓取操作任务。在没有使用我们实验室机器人的任何额外数据,也没有任何任务特定训练或奖励的情况下,该模型成功处理了抓取和拾放等抓取操作任务,包括新物体和新环境。
Planning: We demonstrate that V-JEPA 2-AC, obtained by post-training V-JEPA 2 with only 62 hours of unlabeled robot manipulation data from the popular Droid dataset, can be deployed in new environments to solve prehensile manipulation tasks using planning with given subgoals. Without training on any additional data from robots in our labs, and without any task-specific training or reward, the model successfully handles prehensile manipulation tasks, such as Grasp and Pick-and-Place with novel objects and in new environments.
本文的其余部分组织如下。第 2 节描述了 V-JEPA 2 的预训练过程,包括使其能够超越原始 V-JEPA 配方进行规模扩张的关键要素。第 3 节随后介绍了我们利用预训练的 V-JEPA 2 模型训练任务无关的、以动作条件化的世界模型 V-JEPA 2-AC 的方法。第 4 节演示了通过基于模型的规划使用 V-JEPA 2-AC 进行机器人控制。由于 V-JEPA 2-AC 在学习到的表示空间中建模世界动态,其能力从根本上依赖于 V-JEPA 2 表示空间所捕获的信息,因此我们在第 5 节进一步探讨 V-JEPA 2 在视频理解方面的性能,并在第 6 节探讨其在预测任务中的性能。最后,在第 7 节中,我们展示了 V-JEPA 2 可以与语言模型对齐以进行视频问答。第 8 节讨论相关工作,第 9 节给出结论。
The remainder of this paper is organized as follows. Section 2 describes the V-JEPA 2 pretraining procedure, including the key ingredients enabling scaling beyond the original V-JEPA recipe. Section 3 then introduces our approach to training a task-agnostic action-conditioned world model, V-JEPA 2-AC, leveraging the pretrained V-JEPA 2 model. Section 4 demonstrates using V-JEPA 2-AC for robot control via model-based planning. Because V-JEPA 2-AC models world dynamics in a learned representation space, its capabilities fundamentally depend on the information captured in the V-JEPA 2 representation space, and so we further explore the performance of V-JEPA 2 for video understanding in Section 5 and prediction tasks in Section 6. Finally, in Section 7 we show that V-JEPA 2 can be aligned with a language model for video question answering. Section 8 discusses related work, and we conclude in Section 9.
我们在包含超过 100 万小时视频的视觉数据集上预训练 V-JEPA 2。自监督训练任务基于表示空间中的掩码去噪,并建立在 V-JEPA 框架之上。在本文中,我们通过探索更大规模的模型、增加预训练数据规模,并引入空间和时间渐进式分辨率训练策略来扩展 V-JEPA 框架,从而能够高效地预训练超出短 16 帧视频片段的模型。
We pretrain V-JEPA 2 on a visual dataset that includes over 1 million hours of video. The self-supervised training task is based on mask denoising in representation space and builds upon the V-JEPA framework. In this paper, we extend the V-JEPA framework by exploring larger-scale models, increasing the size of the pretraining data, and introducing a spatial and temporal progressive resolution training strategy that enables us to efficiently pretrain models beyond short 16-frame video clips.
V-JEPA 的目标是从视频的掩码视图 \(x\)(即随机丢弃了部分补丁的视图)预测该视频 \(y\) 的学习表示(图 2 左)。任务元架构由一个编码器 \(E_{\theta}(\cdot)\) 和一个预测器 \(P_{\phi}(\cdot)\) 组成,前者提取视频表示,后者预测掩码视频部分的表示。编码器和预测器通过该目标同时训练。
The V-JEPA objective aims to predict the learned representation of a video \(y\) from a view \(x\) of that video that has been masked, i.e., from which patches have been randomly dropped (Figure 2, left). The task meta-architecture consists of an encoder, \(E_{\theta}(\cdot)\), which extracts video representations, and a predictor, \(P_{\phi}(\cdot)\), which predicts the representation of masked video parts. The encoder and predictor are trained simultaneously using the objective,
其中 \(\Delta_{y}\) 是一个可学习的掩码标记,指示被丢弃补丁的位置。损失函数使用停止梯度操作 \(\text{sg}(\cdot)\) 和编码器网络权重 \(\theta\) 的指数移动平均 \(\overline{\theta}\),以防止表示坍缩。损失仅应用于掩码补丁的预测。
where \(\Delta_{y}\) is a learnable mask token that indicates the locations of the dropped patches. The loss uses a stop-gradient operation, \(\text{sg}(\cdot)\), and an exponential moving average, \(\overline{\theta}\), of the weights \(\theta\) of the encoder network to prevent representation collapse. The loss is applied only to the predictions of the masked patches.
编码器 \(E_{\theta}(\cdot)\) 和预测器 \(P_{\phi}(\cdot)\) 均参数化为视觉 Transformer(ViT)。为了在视觉 Transformer 中编码相对位置信息,我们采用 RoPE(旋转位置编码)而非 [1] 中使用的绝对 sincos 位置编码。我们通过将特征维度划分为三个近似相等的段(分别对应时间、高度和宽度轴),并分别对每个轴应用一维旋转,从而将传统的一维 RoPE 扩展为三维。我们发现,使用 3D-RoPE 而非绝对 sincos 位置嵌入有助于稳定最大模型的训练。为了用我们的 Transformer 编码器处理视频,我们首先将其补丁化为大小为 \(2\times 16\times 16\)(\(T\times H\times W\))的管状序列,并采用与 [1] 相同的多块掩码策略。
The encoder, \(E_{\theta}(\cdot)\), and predictor, \(P_{\phi}(\cdot)\), are each parameterized as a vision transformer (or ViT). To encode relative position information in the vision transformer, we leverage RoPE (Rotary Position Embedding) instead of the absolute sincos position embedding used in [1]. We use a 3D extension of traditional 1D-RoPE by partitioning the feature dimension into three approximately equal segments (for the temporal, height, and width axes) and applying the 1D rotations separately to the segment for each axis. We found that using 3D-RoPE instead of absolute sincos position embeddings helps stabilize training for the largest models. To process a video with our transformer encoder, we first patchify it as a sequence of tubelets of size \(2\times 16\times 16\) ( \(T\times H\times W\) ) and employ the same multiblock masking strategy as in [1].
在本节中,我们引入并研究了四个额外的关键要素,这些要素使得 V-JEPA 预训练原则能够扩展,从而得到我们的 V-JEPA 2 模型。
In this section we introduce and study four additional key ingredients which enable scaling the V-JEPA pre-training principle to obtain our V-JEPA 2 model.
数据扩展:我们通过利用和整理额外的数据源,将数据集规模从 200 万视频增加到 2200 万视频。
Data scaling: We increase the dataset size from 2 million to 22 million videos by leveraging and curating additional data sources.
模型缩放:我们将编码器架构从 3 亿参数扩展到超过 10 亿参数,从 ViT-L 升级到 ViT-g。
Model scaling: We scale the encoder architecture from 300 million to over 1 billion parameters, going from a ViT-L to a ViT-g.
更长的训练:采用 warmup-constant-decay 学习率调度简化了超参数调优,使我们能够将训练从 9 万次迭代延长到 25.2 万次迭代,从而有效利用额外的数据。
Longer training: Adopting a warmup-constant-decay learning rate schedule simplifies hyperparameter tuning and enables us to extend training from 90 thousand up to 252 thousand iterations, effectively leveraging the additional data.
更高分辨率:我们利用 warmup-constant-decay 调度,通过在 warmup 和 constant 阶段训练较短的、较低分辨率的片段,然后在最后的 decay 阶段提高分辨率和/或片段长度,从而高效地扩展到更高分辨率的视频和更长的视频片段。
Higher resolution: We leverage the warmup-constant-decay schedule to efficiently scale to higher resolution video and longer video clips by training on shorter, lower-resolution clips during the warmup and constant phases, and then increasing resolution and/or clip-length during the final decay phase.
本节的其余部分将更详细地描述这些要素中的每一个,并使用接下来描述的评估协议来量化每个要素的影响。
The remainder of this section describes each of these ingredients in further detail and also quantifies the impact of each ingredient using the evaluation protocol described next.
我们模型预训练的目标是将通用的视觉理解注入我们的编码器。因此,我们通过在一组六个运动和外观分类任务上评估模型学习表示的质量来评估我们的模型和数据设计选择:Something-Something v2、Diving-48、Jester、Kinetics、COIN 和 ImageNet。我们使用冻结评估协议:我们冻结编码器权重,并在其表示上训练一个任务特定的 4 层注意力探针来输出预测类别。在本节中,我们主要关注六个理解任务的平均准确率。有关任务、评估协议和结果的更多细节,请参阅第 5 节。
Our goal with model pretraining is to infuse general visual understanding into our encoder. We therefore evaluate our model and data design choices by assessing the quality of the model’s learned representation on a set of six motion and appearance classification tasks: Something-Something v2, Diving-48, Jester, Kinetics, COIN, and ImageNet. We use a frozen evaluation protocol: we freeze the encoder weights and train a task-specific 4-layer attentive probe on its representation to output a predicted class. In this section, we focus mainly on the average accuracy across the six understanding tasks. Refer to Section 5 for additional details about the tasks, evaluation protocol, and results.
我们首先总结缩放分析的关键发现,该分析研究了四个关键因素对下游任务平均性能的影响。图 3 展示了这些缩放干预对 6 个分类任务平均准确率的影响,基线模型为在 200 万视频上使用 V-JEPA 目标预训练的 ViT-L/16 模型。将数据集从 200 万视频扩展到 2200 万视频(VM22M)带来了 1.0 个百分点的提升。将模型从 3 亿参数扩展到 10 亿参数(ViT-g/16)额外带来了 1.5 个百分点的增益。将训练迭代次数从 9 万次扩展到 25.2 万次又贡献了 0.8 个百分点的提升。最后,在预训练和评估阶段同时提高空间分辨率(\(256\rightarrow 384\))和时间长度(\(16\rightarrow 64\)帧),将性能提升至 88.2%,相比 ViT-L/16 基线累计提升了 4.0 个百分点。每项单独的改变都产生了积极影响,证实了视频自监督学习(SSL)中缩放的潜力。
We first present a summary of the key findings of our scaling analysis, where we investigate the impact of the four key ingredients on downstream task average performance. Figure 3 illustrates the effects of these scaling interventions on average accuracy across 6 classification tasks, using a ViT-L/16 model pretrained on 2 million videos with the V-JEPA objective as our baseline. Increasing the dataset from 2 million to 22 million videos (VM22M) yields a 1.0-point improvement. Scaling the model from 300 million to 1 billion parameters (ViT-g/16) provides an additional 1.5-point gain. Extending training from 90K to 252K iterations contributes another 0.8-point improvement. Finally, enhancing both spatial resolution ( \(256\rightarrow 384\) ) and temporal duration ( \(16\rightarrow 64\) frames), during both pretraining and evaluation, boosts performance to 88.2%, representing a cumulative 4.0-point improvement over the ViT-L/16 baseline. Each individual change provides a positive impact, confirming the potential of scaling in video self-supervised learning (SSL).
接下来,我们描述构成预训练数据集的视频和图像来源,以及我们整理数据集的方法。
Next, we describe the sources of videos and images that make up our pretraining dataset, and our approach to curating the dataset.
我们通过结合公开可用的数据源构建了一个大规模视频数据集。使用公开可用的数据源使得其他研究人员能够复现这些结果。整体数据集包括来自 Something-Something v2 数据集(SSv2)的自我中心视频、来自 Kinetics 400、600 和 700 数据集的非自我中心动作视频、来自 HowTo100M 的 YouTube 教程视频,以及来自 YT-Temporal-1B(我们称之为 YT1B)的通用 YouTube 视频。我们还包含来自 ImageNet 数据集的图像,以增加预训练数据的视觉覆盖。为了实现图像和视频的联合预训练,我们在时间上复制图像,并将其视为一个 16 帧的视频,其中所有帧都是相同的。在训练过程中,我们根据经验通过手动调整确定的权重系数从每个数据源进行采样。最终的数据集我们称之为 VideoMix22M(或 VM22M),包含 2200 万个样本。表 1 列出了这些数据源及其权重。
We construct a large-scale video dataset by combining publicly available data sources. Using publicly-available sources in this work enables other researchers to reproduce these results. The overall dataset includes ego-centric videos from the Something-Something v2 dataset (SSv2) introduced in , exo-centric action videos from the Kinetics 400, 600, and 700 datasets , YouTube tutorial videos from HowTo100M , and general YouTube videos from YT-Temporal-1B , which we refer to as YT1B. We also include images from the ImageNet dataset to increase the visual coverage of the pretraining data. To enable joint image and video pretraining, we duplicate an image temporally and treat it as a 16-frame video where all frames are identical. During training, we sample from each data source with a weighting coefficient that we determined empirically via manual tuning. The resulting dataset, which we refer to as VideoMix22M (or VM22M), consists of 22 million samples. Table 1 lists these data sources and their weights.
YT1B 是一个大型视频数据集,包含 140 万小时的视频,与较小的视频数据集(如 Kinetics 和 Something-Something v2)相比,没有进行整理且过滤最少。由于未经整理和不平衡的数据可能阻碍模型性能,我们通过调整现有的基于检索的整理流程来处理视频,从而对 YT1B 进行过滤。具体来说,我们从 YT1B 视频中提取场景,为每个场景计算嵌入向量,然后使用基于聚类的检索过程,根据目标分布选择视频场景,该目标分布由 Kinetics、Something-Something v2、COIN 和 EpicKitchen 训练数据集组成。我们在 A.2 节中描述了数据集构建过程的细节。与类似,我们确保目标验证集中的任何视频都不包含在初始的未整理数据池中。
YT1B is a large video dataset, consisting of 1.4 million video-hours, with no curation and minimal filtering compared to smaller video datasets (like Kinetics and Something-Something v2). Because uncurated and unbalanced data can hinder model performance , we filter YT1B by adapting an existing retrieval-based curation pipeline to handle videos. Specifically, we extract scenes from YT1B videos, compute an embedding vector for each scene, and then use a cluster-based retrieval process to select video scenes according to a target distribution, which is composed of the Kinetics, Something-Something v2, COIN and EpicKitchen training datasets. We describe the details of the dataset construction procedure in Section A.2. Similar to , we ensure that none of the videos from the target validation sets are contained in the initial, uncurated data pool.
在图 4(右)中,我们比较了在未整理的 YT-1B 数据上预训练的 ViT-L 模型与在我们整理的 Curated-YT-1B 数据集上训练的类似模型在视觉理解评估中的平均性能。使用整理后的数据集训练比未整理的基线平均性能提高了+1.4 个百分点。值得注意的是,在 ViT-L 规模下,Curated-YT-1B 训练的模型相对于完整的 VM22M 数据集取得了有竞争力的性能。然而,更大规模的模型从 VM22M 训练中获益更多(见 A.2 节),这表明将 Curated-YT-1B 与其他数据源结合可以增强可扩展性。
In Figure 4 (Right), we compare the average performance on visual understanding evaluations between a ViT-L model pretrained on uncurated YT-1B data and a comparable model trained on our Curated-YT-1B dataset. Training with the curated dataset yields a \(+1.4\) point average performance improvement over the uncurated baseline. Notably, the Curated-YT-1B-trained model achieves competitive performance relative to the full VM22M dataset at the ViT-L scale. However, larger-scale models benefit more from VM22M training (see Section A.2), suggesting that combining Curated-YT-1B with other data sources enhances scalability.
为了探索模型的缩放行为,我们训练了一系列编码器模型,参数数量从 3 亿(ViT-L)到 10 亿(ViT-g)不等。所有编码器架构的细节见附录中的表 12。注意,每个编码器使用相同的预测器架构,类似于 ViT-small。我们在图 5(左)中报告了这些编码器在视觉理解任务上的平均性能。将模型规模从 3 亿(ViT-L)扩展到 10 亿(ViT-g)参数,平均性能提升了+1.5 个百分点。运动和外观理解任务都受益于缩放,SSv2 提升了+1.6 个百分点,Kinetics 提升了+1.5 个百分点(参见表 4)。这些结果证实,自监督视频预训练有效利用了更大的模型容量,直至 10 亿参数的 ViT-g。
To explore the scaling behavior of our model, we trained a family of encoder models with parameter counts ranging from 300 million (ViT-L) to 1 billion (ViT-g) parameters. All encoder architecture details are provided in Table 12 in the appendix. Note that each encoder uses the same predictor architecture, similar to a ViT-small. We report the average performance of these encoders on visual understanding tasks in Figure 5 (Left). Scaling the model size from 300 million (ViT-L) to 1 billion (ViT-g) parameters yields a +1.5 points average performance improvement. Both motion and appearance understanding tasks benefit from scaling, with SSv2 improving by +1.6 points and Kinetics by +1.5 points (cf. Table 4). These results confirm that self-supervised video pretraining effectively leverages larger model capacities, up to the 1B-parameter ViT-g.
V-JEPA 2 模型训练采用预热-恒定学习率调度,随后进入冷却阶段。与[引用]类似,我们发现该调度与半余弦调度性能相当;它还使探索长时间训练运行更具成本效益,因为可以从恒定阶段的多个检查点启动多个冷却运行。我们简化了[引用]中的方案,保持教师 EMA 和权重衰减系数固定,而不使用斜坡调度,因为这些变化对下游理解任务影响极小。图 3 显示,将训练调度从 90K 迭代延长到 252K 迭代,ViT-g 模型的平均性能提升了+0.8 个百分点,验证了延长训练时长的益处。该调度还通过在冷却阶段逐步提高视频分辨率来促进渐进式训练。
V-JEPA 2 model training employs a warmup-constant learning rate schedule followed by a cooldown phase. Similarly to [cite], we found that this schedule performs comparably to a half-cosine schedule; it also makes exploring long training runs more cost-effective, since multiple cooldown runs can be started from different checkpoints of the constant phase. We simplified the recipe from [cite] by maintaining fixed teacher EMA and weight decay coefficients instead of using ramp-up schedule, as these variations showed minimal impact on downstream understanding tasks. Figure 3 shows that extending the training schedule from 90K to 252K iterations yields a +0.8 average performance improvement with ViT-g models, validating the benefits of extended training durations. This schedule also facilitates progressive training by incrementally increasing video resolution during the cooldown phase.
虽然大多数先前的视频编码器专注于 16 帧(约秒)的短视频片段,我们探索了在更高空间分辨率下训练长达 64 帧(16 秒)的更长片段。然而,训练时间随着时长和分辨率的增加而急剧增加——在 64×384×384 输入上训练我们的 ViT-g 模型大约需要 60 个 GPU 年(见图 5,中)。为减少这一开销,我们采用了渐进式分辨率策略,在保持下游性能的同时提高训练效率。我们的训练过程从预热阶段开始,在 12K 次迭代中线性预热学习率,训练 16 帧、256×256 分辨率的视频;随后是主训练阶段,以恒定学习率运行 228K 次迭代。然后,在冷却阶段,我们增加视频时长和分辨率,同时在 12K 次迭代中线性衰减学习率。因此,与在更长时长、更高分辨率视频上训练相关的额外计算开销仅在最终冷却阶段产生。这种方法实现了高效的高分辨率训练:如图 5(中)所示,与在训练的所有阶段从头开始以全分辨率直接训练模型相比,我们实现了 8.4 倍的 GPU 时间减少,同时模型能够处理 64 帧、384×384 分辨率的输入。此外,我们仍然观察到能够处理更长时长和更高分辨率输入的模型的益处,如下文所述。
While most previous video encoders focus on short clips of 16 frames (roughly seconds), we explore training with longer clips of up to 64 frames (16 seconds) at higher spatial resolutions. However, training time increases dramatically with longer durations and higher resolutions — training our ViT-g model on 64×384×384 inputs would require roughly 60 GPU-years (see Figure 5, Middle). To reduce this, we adopt a progressive resolution strategy that boosts training efficiency while maintaining downstream performance. Our training process begins with a warmup phase where we train on 16-frame, 256×256-resolution videos with linear learning rate warmup over 12K iterations, followed by a main training phase with a constant learning rate for 228K iterations. Then, during the cooldown phase, we increase video duration and resolution while linearly decaying the learning rate over 12K iterations. Hence the additional computational overhead associated with training on longer-duration, higher-resolution videos is only incurred during the final cooldown phase. This approach enables efficient high-resolution training: as shown in Figure 5 (Middle), we achieve an 8.4× reduction in GPU time for a model that can ingest 64-frame, 384×384 resolution inputs, compared to directly training such a model from scratch at full resolution throughout all phases of training. Furthermore, we still observe the benefits of a model that can process longer-duration and higher-resolution inputs as discussed next.
预训练之后,V-JEPA 2 模型能够对视频中的缺失部分进行预测。然而,这些预测并未直接考虑智能体可能采取的动作的因果效应。在本节所述的下一训练阶段,我们专注于通过利用少量交互数据使模型对规划有用。为此,我们在冻结的 V-JEPA 2 视频编码器之上学习一个帧因果动作条件预测器(图 2 右侧)。我们使用 Droid 数据集中的数据训练模型,该数据集包含通过遥操作收集的桌面 Franka Panda 机械臂实验数据。我们将得到的动作条件模型称为 V-JEPA 2-AC,并在第 4 节中展示 V-JEPA 2-AC 可用于模型预测控制规划循环中,以在新环境中规划动作。
After pre-training, the V-JEPA 2 model can make predictions about missing parts in videos. However, these predictions do not directly take into account the causal effect of actions that an agent might take. In the next stage of training, described in this section, we focus on making the model useful for planning by leveraging a small amount of interaction data. To that end, we learn a frame-causal action-conditioned predictor on top of the frozen V-JEPA 2 video encoder (Figure 2, right). We train our model on data from the Droid dataset consisting of data from experiments with a table-top Franka Panda robot arm collected through teleoperation. We refer to the resulting action-conditioned model as V-JEPA 2-AC, and in Section 4 we show that V-JEPA 2-AC can be used within a model-predictive control planning loop to plan actions in new environments.
我们的目标是利用预训练后的 V-JEPA 2 模型,获得一个潜在世界模型,该模型可通过闭环模型预测控制用于具身智能体系统的控制。为此,我们训练了 V-JEPA 2-AC,这是一个自回归模型,能够根据控制动作和本体感觉观测来预测未来视频观测的表示。
Our goal is to take the V-JEPA 2 model after pre-training and obtain a latent world model that can be used for control of an embodied agentic system via closed-loop model-predictive control. To achieve this, we train V-JEPA 2-AC, an autoregressive model that predicts representations of future video observations conditioned on control actions and proprioceptive observations.
在本节中,我们针对带有固定外置摄像头的桌面机械臂,描述该框架的具体实例化,其中控制动作对应于末端执行器命令。该模型使用来自原始 Droid 数据集的约 62 小时未标记视频进行训练,该数据集包含短片段(通常为 3-4 秒),展示配备双指夹持器的 7 自由度 Franka Emika Panda 机械臂。这里的未标记视频指的是我们不使用额外的元数据,这些元数据可能指示奖励、每个演示中执行的任务类型,或演示是否成功完成了所尝试的任务。相反,我们仅使用数据集中的原始视频和末端执行器状态信号(数据集中的每个视频都附带元数据,指示每帧的末端执行器状态——三个维度用于位置,三个用于方向,一个用于夹持器状态)。
In this section we describe a concrete instantiation of this framework for a tabletop arm with a fixed exocentric camera, and where control actions correspond to end-effector commands. The model is trained using approximately 62 hours of unlabeled video from the raw Droid dataset, which consists of short videos, typically 3–4 seconds long, of a 7-DoF Franka Emika Panda arm equipped with a two-finger gripper. Here, unlabeled video refers to the fact that we do not use additional meta-data indicating any reward, what type of task was being performed in each demonstration, or whether the demonstration was successful or not in completing the task being attempted. Rather, we only use the raw video and end-effector state signals from the dataset (each video in the dataset is accompanied by meta-data indicating the end-effector state in each frame — three dimensions for position, three for orientation, and one for the gripper state).
在每次训练迭代中,我们从 Droid 数据集中随机采样一个小批量的 4 秒视频片段,为简单起见,丢弃任何短于 4 秒的视频,从而得到数据集的一个较小子集,包含不到 62 小时的视频。视频片段以 \(256\times 256\) 的分辨率和 4 帧/秒(fps)的帧率采样,产生 16 帧的片段,记为 \((x_{k})_{k\in[16]}\) ,其中每个 \(x_{k}\) 代表单个视频帧。每个观测中机器人的末端执行器状态由序列 \((s_{k})_{k\in[16]}\) 表示,其中 \(s_{k}\) 是相对于机器人基座定义的实值 7 维向量。\(s_{k}\) 的前三个维度编码末端执行器的笛卡尔位置,接下来三个维度以外在欧拉角形式编码其方向,最后一个维度编码夹持器状态。我们通过计算相邻帧之间末端执行器状态的变化来构造动作序列 \((a_{k})_{k\in[15]}\) 。具体来说,每个动作 \(a_{k}\) 是一个实值 7 维向量,表示帧 \(k\) 和 \(k+1\) 之间末端执行器状态的变化。我们对采样的视频片段应用随机调整大小裁剪增强,宽高比在 (0.75, 1.35) 范围内采样。
In each iteration of training we randomly sample a mini-batch of 4 second video clips from the Droid dataset, and, for simplicity, discard any videos shorter than 4 seconds, leaving us with a smaller subset of the dataset comprising under 62 hours of video. The video clips are sampled with resolution \(256\times 256\) and a frame-rate of 4 frames-per-second (fps), yielding 16 frame clips denoted by \((x_{k})_{k\in[16]}\) , where each \(x_{k}\) represents a single video frame. The robot’s end-effector state in each observation is denoted by the sequence \((s_{k})_{k\in[16]}\) , where \(s_{k}\) is a real-valued 7D vector defined relative to the base of the robot. The first three dimensions of \(s_{k}\) encode the cartesian position of the end-effector, the next three dimensions encode its orientation in the form of extrinsic Euler angles, and the last dimension encodes the gripper state. We construct a sequence of actions \((a_{k})_{k\in[15]}\) by computing the change in end-effector state between adjacent frames. Specifically, each action \(a_{k}\) is a real-valued 7-dimensional vector representing the change in end-effector state between frames \(k\) and \(k+1\) . We apply random-resize-crop augmentations to the sampled video clips with the aspect-ratio sampled in the range (0.75, 1.35).
我们使用 V-JEPA 2 编码器 \(E(\cdot)\) 作为图像编码器,独立编码给定片段中的每一帧,以获得特征图序列 \((z_{k})_{k\in[16]}\) ,其中 \(z_{k}\coloneqq E(x_{k})\in\mathbb{R}^{H\times W\times D}\) ,\(H\times W\) 表示特征图的空间分辨率,\(D\) 表示嵌入维度。实际上,我们的特征图使用 ViT-g 编码器编码,形状为 \(16\times 16\times 1408\) 。注意,在后训练阶段,编码器保持冻结。特征图序列、末端执行器状态和动作在时间上交错为 \((a_{k},s_{k},z_{k})_{k\in[15]}\) ,并通过 Transformer 预测网络 \(P_{\phi}(\cdot)\) 处理,以获得下一状态表示预测序列 \((\hat{z}_{k+1})_{k\in[15]}\) 。标量值的教师强制损失函数最终计算为
We use V-JEPA 2 encoder \(E(\cdot)\) as an image encoder and encode each frame independently in a given clip to obtain a sequence of feature maps \((z_{k})_{k\in[16]}\) , where \(z_{k}\coloneqq E(x_{k})\in\mathbb{R}^{H\times W\times D}\) with \(H\times W\) denoting the spatial resolution of the feature map, and \(D\) the embedding dimension. In practice, our feature maps are encoded using the ViT-g encoder and have the shape \(16\times 16\times 1408\) . Note that the encoder is kept frozen during this post-training phase. The sequence of feature maps, end-effector states, and actions are temporally interleaved as \((a_{k},s_{k},z_{k})_{k\in[15]}\) and processed with the transformer predictor network \(P_{\phi}(\cdot)\) to obtain a sequence of next state representation predictions \((\hat{z}_{k+1})_{k\in[15]}\) . The scalar-valued teacher-forcing loss function is finally computed as
其中 \(T=15\) 。我们还计算两步展开损失,以提高模型在推理时进行自回归展开的能力。为简化说明并略微重载符号,令 \(P_{\phi}(\hat{a}_{1:T};s_{k},z_{k})\in\mathbb{R}^{H\times W\times D}\) 表示通过自回归运行 V-JEPA 2-AC 并给定动作序列 \((\hat{a}_{i})_{i\in[T]}\) 从 ( \(s_{k}\) , \(z_{k}\) ) 开始获得的最终预测状态表示。我们现在可以将展开损失表示为:
with \(T=15\) . We also compute a two-step rollout loss to improve the model’s ability to perform autoregressive rollouts at inference time. For simplicity of exposition and with slight overloading of notation, let \(P_{\phi}(\hat{a}_{1:T};s_{k},z_{k})\in\mathbb{R}^{H\times W\times D}\) denote the final predicted state representation obtained by autoregressively running V-JEPA 2-AC with an action sequence \((\hat{a}_{i})_{i\in[T]}\) , starting from ( \(s_{k}\) , \(z_{k}\) ). We can now denote the rollout loss as:
在实践中,我们使用 \(T=2\) 来计算 rollout 损失,这样我们只通过一个循环步骤对预测器进行微分。
In practice we use \(T=2\) for computing the rollout loss, such that we only differentiate the predictor through one recurrent step.
因此,整体训练目标由下式给出:
The overall training objective is thus given by
并且相对于预测器权重 \(\phi\) 进行最小化。为了说明目的,训练过程在图 6 中示出,其中 \(T=4\) 同时用于教师强制和 rollout 损失。
and is minimized with respect to the predictor weights \(\phi\) . For illustrative purposes, the training procedure is depicted in Figure 6 with \(T=4\) for both the teacher forcing and rollout loss.
预测器网络 \(P_{\phi}(\cdot)\) 是一个约 3 亿参数的 Transformer 网络,具有 24 层、16 个头、1024 隐藏维度和 GELU 激活。输入到预测器的动作、末端执行器状态和展平的特征图通过单独的可学习仿射变换映射到预测器的隐藏维度。类似地,预测器最后一个注意力块的输出经过可学习仿射变换,映射回编码器的嵌入维度。我们使用我们的 3D-RoPE 实现来表示展平特征图中每个视频补丁的时空位置,而仅对动作和姿态标记应用时间旋转位置嵌入。我们在预测器中使用块因果注意力模式,以便给定时间步的每个补丁特征可以关注同一时间步的动作、末端执行器状态和其他补丁特征,以及之前时间步的特征。
The predictor network \(P_{\phi}(\cdot)\) is a \(\sim\) 300M parameter transformer network with 24-layers, 16 heads, 1024 hidden dimension, and GELU activations. The action, end-effector state, and flattened feature maps input to the predictor are processed with separate learnable affine transformations to map them to the hidden dimension of the predictor. Similarly, the outputs of the last attention block of the predictor go through a learnable affine transformation to map them back to the embedding dimension of the encoder. We use our 3D-RoPE implementation to represent the spatiotemporal position of each video patch in the flattened feature map, while only applying the temporal rotary positional embeddings to the action and pose tokens. We use a block-causal attention pattern in the predictor so that each patch feature at a given time step can attend to the action, end-effector state, and other patch features from the same timestep, as well as those from previous time steps.
给定目标状态的图像,我们利用 V-JEPA 2-AC 通过规划来执行下游任务。具体而言,在每个时间步,我们通过最小化目标条件能量函数来规划固定时间范围内的动作序列。然后执行第一个动作,观察新状态,并重复该过程。令 \(s_{k}\) 表示当前末端执行器状态,\(x_{k}\) 和 \(x_{g}\) 分别表示当前观测帧和目标图像,它们分别通过视频编码器编码得到特征图 \(z_{k}\) 和 \(z_{g}\) 。给定规划范围 \(T\) ,我们通过最小化目标条件能量函数来优化机器人动作序列 \((a^{\star}_{i})_{i\in[T]}\) ,
Given an image of the goal state, we leverage V-JEPA 2-AC for downstream tasks by planning. Specifically, at each time step, we plan an action sequence for a fixed time horizon by minimizing a goal-conditioned energy function. We then execute the first action, observe the new state, and repeat the process. Let \(s_{k}\) denote the current end-effector state, and \(x_{k}\) and \(x_{g}\) denote the current observed frame and goal image, respectively, which are separately encoded with the video encoder to obtain the feature maps \(z_{k}\) and \(z_{g}\) . Given a planning horizon, \(T\) , we optimize a sequence of robot actions, \((a^{\star}_{i})_{i\in[T]}\) , by minimizing a goal-conditioned energy function,
即 \((a^{\star}_{i})_{i\in[T]}\coloneqq\text{argmin}_{\hat{a}_{1:T}}\ \mathcal{E}(\hat{a}_{1:T};\ z_{k},s_{k},z_{g})\) 。如图 7 所示,模型通过选择一条轨迹来推断动作序列 \((a^{\star}_{i})_{i\in[T]}\) ,该轨迹最小化世界模型在未来 \(T\) 步的想象状态表示与其目标表示之间的 L1 距离。在实践中,我们在每个规划步骤中使用交叉熵方法最小化 (5),并且只执行第一个动作,然后重新规划,如同滚动时域控制。
such that \((a^{\star}_{i})_{i\in[T]}\coloneqq\text{argmin}_{\hat{a}_{1:T}}\ \mathcal{E}(\hat{a}_{1:T};\ z_{k},s_{k},z_{g})\) . As illustrated in Figure 7, the model infers an action sequence \((a^{\star}_{i})_{i\in[T]}\) by selecting a trajectory that minimizes the L1 distance between the world model’s imagined state representation \(T\) steps into the future and its goal representation. In practice, we minimize (5) in each planning step using the Cross-Entropy Method , and only execute the first action on the robot before re-planning, as in receding horizon control.
在本节中,我们演示如何使用 V-JEPA 2-AC 通过模型预测控制来实现基本的机器人技能,如到达、抓取和拾取放置。我们专注于具有视觉目标指定的任务,并展示 V-JEPA 2-AC 能够零样本泛化到新环境。
In this section we demonstrate how V-JEPA 2-AC can be used to implement basic robot skills like reaching, grasping, and pick-and-place via model-predictive control. We focus on tasks with visual goal specification and show that V-JEPA 2-AC generalizes zero-shot to new environments.
我们将 V-JEPA 2-AC 与两个基线模型进行比较:一个是通过行为克隆训练的视觉-语言-动作模型,另一个是基于视频生成的世界模型。
We compare the performance of V-JEPA 2-AC with two baselines: one vision-language-action model trained with behavior cloning, and one video generation-based world model.
第一个基线基于 Octo 视频-语言-动作模型,该模型支持目标图像条件。我们从该模型的 octo-base-1.5 版本的开源权重开始,该版本在包含超过 100 万条轨迹的 Open-X Embodiment 数据集上进行了预训练。相比之下,我们在 Droid 数据集的 23k 条轨迹上训练 V-JEPA 2-AC,包括成功和失败的轨迹。我们使用行为克隆在完整的 Droid 数据集上微调 Octo 模型,并使用图像目标和末端执行器状态进行事后重标记。具体来说,在训练期间,我们从 Droid 数据集中随机采样轨迹片段,并均匀采样轨迹中向前最多 20 个时间步的目标图像。我们使用官方开源代码进行微调,包括所有标准的 Droid 优化超参数,并利用单侧图像视图输入,分辨率为 \(256\times 256\),上下文为前两帧,未来动作范围为 4 个动作。
The first baseline is based on the Octo video-language-action model that allows for goal-image conditioning. We start from the open-source weights of the octo-base-1.5 version of the model, which is pretrained on the Open-X Embodiment dataset containing over 1M trajectories. In comparison, we train V-JEPA 2-AC on 23k trajectories from Droid, including successes and failures. We fine-tune the Octo model with behavior cloning on the entire Droid dataset using hindsight relabeling with image goals and end-effector states. In particular, we sample random segments of trajectories from the Droid dataset during training, and uniformly sample goal images up to 20 timesteps forward in the trajectory. We use the official open-source code for fine-tuning, including all standard Droid optimization hyperparameters, and leverage single side image view inputs at \(256\times 256\) resolution, a context of two previous frames, and a horizon of 4 future actions.
我们比较的第二个基线基于 Cosmos 视频生成模型。我们从无动作 Cosmos 模型(带有连续分词器的潜在扩散 7B 模型)的开源权重开始,该模型在 2000 万小时的视频上进行了训练,我们使用官方发布的动作条件微调代码在 Droid 上对模型进行微调。为了提高在 Droid 上训练的性能,我们 (i) 降低了学习率,以匹配视频条件 Cosmos 配方中使用的学习率;(ii) 移除了视频条件中的 dropout,以改善训练动态;(iii) 将噪声水平提高了 \(e^{2}\) 倍,因为我们观察到,使用较低噪声因子训练的模型难以利用条件帧中的信息。尽管 Cosmos 技术报告提到将世界模型用于规划或模型预测控制作为未来应用,但据我们所知,这是首次报道使用 Cosmos 模型进行机器人控制的尝试。
The second baseline we compare with is based on the Cosmos video generation model. We start with the open-source weights for the action-free Cosmos model (latent diffusion-7B with continuous tokenizer), which was trained on 20M hours of video, and we fine-tune the model on Droid using the officially-released action-conditioned fine-tuning code. To improve performance when training on Droid, we (i) lowered the learning rate to match that used in the video-conditioned Cosmos recipe, (ii) removed the dropout in the video conditioning to improve the training dynamics, and (iii) increased the noise level by a factor of \(e^{2}\), as we observed that the model trained with a lower noise factor struggled to leverage the information in the conditioning frame. Although the Cosmos technical report mentions using world models for planning or model-predictive control as a future application, to the best of our knowledge this is the first reported attempt using Cosmos models for robot control.
所有模型均在两个不同实验室的 Franka Emika Panda 机械臂(配备 RobotiQ 夹爪)上进行零样本部署,这两个实验室均未出现在 Droid 数据集中。视觉输入通过未校准的低分辨率单目 RGB 相机提供。机器人使用完全相同的模型权重和推理代码,以及基于操作空间控制的类似低级控制器。我们对 V-JEPA 2-AC 世界模型和 Cosmos 世界模型使用阻塞控制(即系统等待最后一个命令动作完成后再向控制器发送新动作),并对 Octo 尝试了阻塞和非阻塞控制,并报告两种选项中的最佳性能。当使用 V-JEPA 2-AC 和 Cosmos 进行规划时,我们将每个采样动作限制在以原点为中心、半径为 \(0.075\) 的 L1 球内,这对应于每个单独动作的最大末端执行器位移约为 13 厘米,因为较大的动作相对而言超出模型的分布范围。
All models are deployed zero-shot on Franka Emika Panda arms with RobotiQ grippers, located in two different labs, neither of which appears in the Droid dataset. Visual input is provided through an uncalibrated low-resolution monocular RGB camera. The robots use the same exact model weights and inference code, and similar low-level controllers based on operational space control. We use blocking control for both the V-JEPA 2-AC world model and Cosmos world model (i.e., the system waits for the last commanded action to be completed before sending a new action to the controller) and experiment with both blocking and non-blocking control for Octo, and report the best performance across the two options. When planning with V-JEPA 2-AC and Cosmos, we constrain each sampled action to the L1-Ball of radius \(0.075\) centered at the origin, which corresponds to a maximum end-effector displacement of approximately 13 cm for each individual action, since large actions are relatively out-of-distribution for the models.
首先,我们在单目标到达任务上进行评估,该任务要求根据单个目标图像将末端执行器移动到空间中的期望位置。该任务衡量对动作的基本理解以及从单目 RGB 相机对场景(包括深度)的 3D 空间理解。
First, we evaluate on the task of single-goal reaching, which involves moving the end-effector to a desired location in space based on a single goal image. This task measures for a basic understanding of actions as well as a 3D spatial understanding of the scene (including depth) from the monocular RGB camera.
在图 9 中,我们可视化了方程(5)中 V-JEPA 2-AC 的能量景观,针对\(\Delta y\)到达任务,作为单个笛卡尔控制动作的函数,扫描\(\Delta x\)和\(\Delta y\),同时保持\(\Delta z=0\)固定。能量函数在接近真实动作处取得最小值,进一步证明模型已学会合理推断动作的效果,而无需精确传感。同样有趣的是,V-JEPA 2-AC 诱导的能量景观相对平滑且局部凸,这应有助于规划。
In Figure 9, we visualize the V-JEPA 2-AC energy landscape from equation (5) for the \(\Delta y\) reaching task as a function of a single cartesian-control action, sweeping \(\Delta x\) and \(\Delta y\) while holding \(\Delta z=0\) fixed. The energy function achieves its minimum near the ground-truth action, providing further evidence that the model has learned to reasonably infer the effect of actions without requiring precision sensing. It is also interesting to observe that the energy landscape induced by V-JEPA 2-AC is relatively smooth and locally convex, which should facilitate planning.
接下来,我们在更具挑战性的抓取物体操作任务上评估所有模型,即抓取、带物体到达和拾放。成功率报告在表 2 和表 3 中,并在 10 次试验中取平均,每次试验任务有各种排列(例如,物体位置、起始姿态等)。对于抓取和带物体到达任务,模型只显示单个目标图像。对于拾放任务,除了最终目标外,我们还向模型展示两个子目标图像。第一个目标图像显示物体被抓取,第二个目标图像显示物体在目标位置附近。模型首先针对第一个子目标优化动作 4 个时间步,然后自动切换到第二个子目标进行接下来的 10 个时间步,最后针对第三个目标进行最后 4 个时间步。拾放任务的机器人执行示例见图 10。实验室 1 中所有单独任务的起始和目标帧见 B.2 节。抓取任务需要从视觉反馈中进行精确控制以正确抓取物体。带物体到达任务要求模型在持有物体时导航,这需要基本的直观物理理解以避免掉落物体。最后,拾放任务测试组合这些原子技能的能力。
Next, we evaluate all models on more challenging prehensile object manipulation tasks, namely grasp, reach with object, and pick-and-place. Success rates are reported in Table 2 and Table 3, and averaged across 10 trials with various permutations to the task across trials (e.g., object location, starting pose, etc.). For the grasp and reach with object tasks the model is shown a single goal image. For the pick-and-place tasks we present two sub-goal images to the model in addition to the final goal. The first goal image shows the object being grasped, the second goal image shows the object in the vicinity of the goal position. The model first optimizes actions with respect to the first sub-goal for 4 time-steps before automatically switching to the second sub-goal for the next 10 time-steps, and finally the third goal for the last 4 time-steps. Examples of robot execution for the pick-and-place task are shown in Figure 10. Start and goal frames for all individual tasks in Lab 1 are shown in Section B.2. The grasp task requires precise control from visual feedback to correctly grip the object. The reach with object task requires the model to navigate while holding an object, which necessitates a basic understanding of intuitive physics to avoid dropping the object. Finally, the pick-and-place task tests for the ability to compose these atomic skills.
虽然所有模型在到达任务上都取得了高成功率,但在涉及物体交互的任务上性能差异更为明显。我们观察到所有模型的成功率取决于被操作物体的类型。例如,我们发现杯子最容易通过将一个手指放入物体内部并抓住边缘来抓取,但如果模型产生的控制动作不够精确,机器人将错过杯子的边缘而无法抓取物体。操作盒子时,有更多可行的抓取配置,但模型需要更精确的夹持器控制以确保手指张开足够宽以抓取物体。我们看到,对于所有模型,成功率随物体类型的变化是由于次优动作和操作各自物体所特有的挑战的组合。尽管如此,我们看到 V-JEPA 2-AC 模型在所有任务上取得了最高的成功率,突出了潜在规划在机器人操作中的可行性。
While all models achieve a high success-rate on reach, differences in performance are more apparent on tasks involving object interaction. We observe that the success-rate for all models depends on the type of object being manipulated. For instance, we find that the cup is mostly easily grasped by placing one finger inside the object and gripping around the rim, however if the control actions produced by the model are not accurate enough, the robot will miss the rim of the cup and fail to grasp the object. When manipulating the box, there are many more feasible grasping configurations, however, the model requires more precise gripper control to ensure that the fingers are open wide enough to grasp the object. We see that, for all models, the variation in success-rate with respect to the object type is due to the combination of sub-optimal actions and the unique challenges associated with manipulating each respective object. Nonetheless, we see that the V-JEPA 2-AC model achieves the highest success-rate across all tasks, highlighting the feasibility of latent planning for robot manipulation.
在表 3 中,我们比较了使用 V-JEPA 2-AC 与基于潜在扩散的 Cosmos 动作条件视频生成模型时的规划性能。在两种情况下,我们都利用交叉熵方法优化动作序列,使用单个 NVIDIA RTX 4090 GPU,并通过在模型的潜在空间中编码目标帧来构建能量函数,如方程(5)所示。使用 80 个样本、10 个细化步骤和规划视野为 1 时,Cosmos 在每一步规划中计算单个动作需要 4 分钟。虽然使用 Cosmos 在到达任务上取得了 80%的高成功率,但在物体交互任务上性能较弱。注意,在每动作 4 分钟的规划时间下,完整的拾放轨迹需要超过一小时的机器人执行时间。相比之下,每个细化步骤使用 10 倍多的样本,V-JEPA 2-AC 世界模型每个动作仅需 16 秒,并在所有考虑的机器人技能上取得更高性能。我们可以在未来工作中通过利用额外的计算资源进行规划、减少每个时间步使用的样本和细化步骤数量、在世界模型的想象中训练前馈策略以初始化规划问题,或可能利用基于梯度的规划(对于 V-JEPA 2-AC)来潜在地减少两种模型的规划时间。
In Table 3, we compare planning performance when using V-JEPA 2-AC versus the Cosmos action-conditioned video generation model based on latent diffusion. In both cases we leverage the cross-entropy method for optimizing the sequence of actions using a single NVIDIA RTX 4090 GPU, and construct the energy function by encoding the goal frame in the latent space of the model, as in equation (5). With 80 samples, 10 refinement steps, and a planning horizon of 1, it takes 4 minutes to compute a single action in each planning step with Cosmos. While we achieve a high success rate of 80% on the reach tasks when using Cosmos, performance on object interaction tasks is weaker. Note that under a planning time of 4 minutes per action, a full pick & place trajectory requires over one hour of robot execution. By contrast, with 10 \(\times\) more samples in each refinement step, the V-JEPA 2-AC world model requires only 16 seconds per action and leads to higher performance across all considered robot skills. We can potentially reduce the planning time for both models in future work by leveraging additional computing resources for planning, reducing the number of samples and refinement steps used at each time step, training a feed-froward policy in the world-models’ imagination to initialize the planning problem, or potentially leveraging gradient-based planning in the case of V-JEPA 2-AC.
由于 V-JEPA 2-AC 模型在给定末端执行器笛卡尔控制动作的情况下,被训练用于预测下一视频帧的表示,且没有显式的相机标定,因此它必须从单目 RGB 相机输入中隐式推断动作坐标轴。然而,在许多情况下,机器人基座在相机视野中不可见,因此推断动作坐标轴的问题定义不明确,导致世界模型出现误差。在实践中,我们手动尝试了不同的相机位置,最终选定了一个在所有实验中均表现良好的位置。我们在第 B.4 节中对 V-JEPA 2-AC 世界模型对相机位置的敏感性进行了定量分析。
Since the V-JEPA 2-AC model is trained to predict representations of the next video frame given an end-effector Cartesian control action, without any explicit camera calibration, it must therefore implicitly infer the action coordinate axis from the monocular RGB camera input. However, in many cases, the robot base is not visible in the camera frame, and thus the problem of inferring the action coordinate axis is not well defined, leading to errors in the world model. In practice, we manually tried different camera positions before settling on one that worked well across all of our experiments. We conduct a quantitative analysis of the V-JEPA 2-AC world model’s sensitivity to camera position in Section B.4.
使用世界模型进行长时程规划受到多种因素的限制。首先,自回归预测存在误差累积问题:表示空间预测的准确性随着自回归展开长度的增加而下降,从而使得在长时程上可靠规划变得更加困难。其次,长时程规划增加了搜索空间的大小:随着规划时域线性增加,可能的动作轨迹数量呈指数增长,这使得在计算上难以进行长时程规划。另一方面,长时程规划对于解决非贪婪预测任务(例如,没有图像子目标的无序抓取放置)是必要的。未来探索用于长时程规划的世界模型的工作将有助于解决更多复杂且有趣的任务。
Long horizon planning with world models is limited by a number of factors. First, autoregressive prediction suffers from error accumulation: the accuracy of the representation-space predictions decreases with longer autoregressive rollouts, thereby making it more difficult to reliably plan over long horizons. Second, long-horizon planning increases the size of the search space: the number of possible action trajectories increases exponentially given a linear increase in the planning horizon, thereby making it computationally challenging to plan over long horizons. On the other hand, long-horizon planning is necessary for solving non-greedy prediction tasks, e.g., pick-and-place without image sub-goals. Future work exploring world models for long-horizon planning will enable the solution of many more complex and interesting tasks.
遵循许多先前在目标条件机器人操作方面的工作,我们当前的优化目标公式假设我们可以访问视觉目标。然而,在野外部署机器人时,用其他形式(如语言)来表达目标可能更为自然。未来将潜在的动作条件世界模型与语言模型对齐的工作,将朝着通过自然语言进行更通用的任务规范迈进。
Following many previous works in goal-conditioned robot manipulation, our current formulation of the optimization target assumes that we have access to visual goals. However, when deploying robots in-the-wild, it may be more natural to express goals in other forms, such as with language. Future work that aligns latent action-conditioned world models with language models will step towards more general task specification via natural language.
表示空间世界模型(如上述 V-JEPA 2-AC)的能力本质上受限于学习到的表示空间中所编码的状态信息。在本节及后续章节中,我们探究 V-JEPA 2 学习到的表示,并将 V-JEPA 2 编码器与其他视觉编码器在视觉分类任务上进行对比。
The capabilities of a representation-space world model, such as V-JEPA 2-AC discussed above, are inherently limited by the state information encoded in the learned representation space. In this section and subsequent sections, we probe the representations learned by V-JEPA 2 and compare the V-JEPA 2 encoder to other vision encoders on visual classification.
视觉分类任务可侧重于外观理解或运动理解。外观理解任务通常可利用输入视频片段中单帧可见的信息来解决(即使分类标签描述的是动作),而运动理解任务则需要多帧才能正确分类视频。为确保对运动和外观的均衡评估,我们选取了三个运动理解任务:Something-Something v2 (SSv2)、Diving-48 和 Jester,这些任务要求模型理解人类手势和动作。对于外观理解,我们选择了 Kinetics400 (K400)、COIN 和 ImageNet (IN1K),这些任务涉及动作、场景和物体的识别。实验表明,V-JEPA 2 在运动理解任务上优于最先进的视觉编码器,同时在外观理解任务上具有竞争力。
Visual classification tasks can focus either on appearance understanding or motion understanding. While appearance understanding tasks can generally be solved using information visible in a single frame of an input video clip (even when the classification labels describe actions), motion understanding tasks require several frames to correctly classify a video. To ensure a balanced evaluation of both motion and appearance, we have selected three motion understanding tasks, namely Something-Something v2 (SSv2), Diving-48, and Jester, which require the model to understand human gestures and movements. For appearance understanding, we have chosen Kinetics400 (K400), COIN, and ImageNet (IN1K), which involve recognizing actions, scenes, and objects. Empirically, we show that V-JEPA 2 outperforms state-of-the-art visual encoders on motion understanding tasks, while being competitive on appearance understanding tasks.
我们使用每个任务的训练数据,在冻结的编码器输出之上训练一个 4 层注意力探针。我们的注意力探针由四个 Transformer 块组成,最后一个块将标准自注意力替换为使用可学习查询令牌的交叉注意力层。按照标准做法,在推理过程中从视频中采样多个固定帧数的片段,然后将分类 logits 在片段间取平均。我们保持与 V-JEPA 2 预训练时相似的分辨率。我们在附录 C.2 中对注意力探针的层数进行了消融实验,并提供了下游任务中使用的片段数量、片段大小及其他超参数的完整细节。
We train a 4-layer attentive probe on top of the frozen encoder output using the training data from each task. Our attentive probe is composed of four transformer blocks, the last of which replaces standard self-attention with a cross-attention layer using a learnable query token. Following standard practice, several clips with a fixed number of frames are sampled from a video during inference. The classification logits are then averaged across clips. We keep the resolution similar to the one used for V-JEPA 2 pretraining. We ablate the number of layers of our attentive probe in Section C.2, and also provide full details on the number of clips, clip size, and other hyperparameters used in the downstream tasks.
我们将 V-JEPA 2 在运动和外观任务上的性能与其他几种视觉编码器进行比较:带寄存器的 DINOv2 是当前图像自监督学习的最先进模型,而 SigLIP2 和感知编码器 PEcoreG 是图像-文本对比预训练的两个最先进模型。我们还考虑了两种视频编码器:自监督的 V-JEPA,以及主要依赖视觉-文本对比预训练的 InternVideo2s2-1B。
We compare the performance of V-JEPA 2 on motion and appearance tasks with several other visual encoders: DINOv2 with registers is the current state-of-the-art model for self-supervised learning with images, while SigLIP2 and the Perception Encoder PEcoreG are two state-of-the-art models for image-text contrastive pretraining. We also consider two video encoders: the self-supervised V-JEPA, and InternVideo2s2-1B which relies primarily on vision-text contrastive pretraining.
我们对每个基线和 V-JEPA 2 使用相同的评估协议,在冻结的编码器之上学习一个注意力探针,类似于 [参考文献]。我们按照 [参考文献] 中的过程将基于图像的模型适应到视频,拼接每个输入帧的特征。对于 InternVideo2s2-1B,我们在 ImageNet 任务中使用其图像位置嵌入,对于视频任务,我们将其位置嵌入从 4 帧插值到 8 帧,产生与 V-JEPA 2 相似的令牌数量。尽管使用共同的评估协议,基线编码器在不同数据上训练(例如,DINOv2 在 LVD-142M 上,PEcoreG 在 MetaCLIP 上),因此不能直接比较。因此,我们只能在系统层面比较不同方法;即,尽管训练协议和数据不同,但使用一致的评估协议。我们还纳入了文献中采用类似冻结协议但可能使用不同注意力头架构的现有结果。特别是,我们分享了 VideoMAEv2、InternVideo-1B 和 6B 以及 VideoPrism 在我们所考虑的分类任务上的报告结果(如有)。我们在附录 C.1 中提供了完整的评估和超参数。
We use the same evaluation protocol for every baseline and for V-JEPA 2, learning an attentive probe on top of the frozen encoder, similar to [reference]. We adapt image-based models to video following the procedure used in [reference], concatenating the features of each input frame. For InternVideo2s2-1B, we use its image positional embedding for the ImageNet task, and for video tasks we interpolate its positional embedding from 4 frames to 8, producing a token count similar to V-JEPA 2. Despite using a common evaluation protocol, the baseline encoders are trained on different data (e.g., DINOv2 on LVD-142M, PEcoreG on MetaCLIP) and are thus not directly comparable. We can therefore only compare different approaches at a system level; i.e., with a consistent evaluation protocol despite differences in training protocol and data. We also include existing results from the literature using a similar frozen protocol, but with potentially different attentive head architecture. In particular, we share reported results of VideoMAEv2, InternVideo-1B and 6B, and VideoPrism on the classification tasks we consider, when available. We provide complete evaluation and hyperparameters in Section C.1.
动作预测(action anticipation)是指根据动作发生前一段时间的上下文视频片段来预测未来的动作。利用 Epic-Kitchens-100(EK100)基准,我们证明了 V-JEPA 2 的动作预测性能随模型规模增大而持续提升。此外,尽管仅使用了在 V-JEPA 2 表示之上训练的注意力探针,我们表明 V-JEPA 2 显著优于先前专门为这一任务设计的最先进方法。
Action anticipation consists in predicting the future action given a contextual video clip leading up to some time before the action. Using the Epic-Kitchens-100 (EK100) benchmark, we demonstrate that V-JEPA 2 action anticipation performance increases consistently with model size. Furthermore, despite only using an attentive probe trained on top of V-JEPA 2 representations, we show that V-JEPA 2 significantly outperforms prior state-of-the-art approaches that were specifically designed for this task.
EK100 数据集包含从 45 个厨房环境中以自我中心视角录制的 100 小时烹饪活动。EK100 中的每个视频都标注了动作片段,包括开始时间戳、结束时间戳和动作标签。共有 3,568 个独特的动作标签,每个标签由动词和名词类别组成,总共有 97 个动词类别和 300 个名词类别。EK100 动作预测任务涉及从动作片段开始时间戳之前的视频片段(称为上下文)预测名词、动词和动作(即联合预测动词和名词)。上下文结束与动作片段开始之间的间隔为预测时间,默认设置为 1 秒。鉴于给定上下文可能对应不同的未来动作,使用平均类别召回率@5(mean-class recall-at-5)作为性能度量指标。
The EK100 dataset is comprised of 100 hours of cooking activities recorded from an egocentric perspective across 45 kitchen environments. Each video in EK100 is annotated with action segments, which include a start timestamp, an end timestamp, and an action label. There are 3,568 unique action labels, each consisting of a verb and a noun category, with a total of 97 verb categories and 300 noun categories. The EK100 action anticipation task involves predicting noun, verb, and action (i.e., predicting verb and noun jointly) from a video clip, referred to as context, that occurs before the start timestamp of an action segment. The interval between the end of the context and the beginning of the action segment is the anticipation time, which is set to 1 second by default. Given that different future actions may be possible from a given context, mean-class recall-at-5 is used as the metric to measure performance.
我们在冻结的 V-JEPA 2 编码器和预测器之上训练一个注意力探针来预测未来动作。具体来说,我们采样一个在动作开始前 1 秒结束的视频片段。该视频上下文被输入 V-JEPA 2 编码器。预测器接收编码器表示以及对应于未来 1 秒帧的掩码标记,并预测未来视频帧的表示。预测器和编码器的输出在标记维度上拼接,并输入到与第 5 节所用架构类似的注意力探针中,区别在于预测探针的最终交叉注意力层学习三个查询标记(而不是一个),每个查询输出分别输入到不同的线性分类器,以分别预测动作类别、动词类别和名词类别。对每个分类器独立应用焦点损失(focal loss),然后在通过探针的共享注意力块进行反向传播之前求和。我们在第 D.1 节中提供了额外的细节和评估超参数。
An attentive probe is trained on top of the frozen V-JEPA 2 encoder and predictor to anticipate future actions. Specifically, we sample a video clip that ends 1 second before an action starts. This video context is fed to the V-JEPA 2 encoder. The predictor takes the encoder representation, along with the mask tokens corresponding to the frame 1 second into the future, and predicts the representation of the future video frame. The outputs of the predictor and encoder are concatenated along the token dimension and fed to an attentive probe with a similar architecture to those used in Section 5, with the difference being that the anticipation probe’s final cross-attention layer learns three query tokens (as opposed to one), and each query output is fed to a different linear classifier to predict the action category, the verb category, and the noun category respectively. A focal loss is applied to each classifier independently and then summed before backpropagating through the shared attention blocks of the probe. We provide additional details and evaluation hyperparameters in Section D.1.
我们将我们的模型与三个专门为动作预测训练的基线进行比较:InAViT 是一种利用显式手-物交互建模的监督方法,而 Video-LLaMA 和 PlausiVL 都是利用大型语言模型的方法,参数规模高达 70 亿。
We compare our model with three baselines that are trained specifically for action anticipation: InAViT is a supervised approach that leverages explicit hand-object interaction modeling, and Video-LLaMA and PlausiVL are both approaches that leverage a large language model, with up to 7 billion parameters.
V-JEPA 2 显著优于之前的最先进模型 PlausiVL,即使其参数量为 3 亿,而 PlausiVL 使用了 80 亿参数。特别是,V-JEPA 2 ViT-g384 在动作召回率@5 上比 PlausiVL 提高了 \(+12.1\) 个百分点,相当于相对改进 \(44\%\)。
V-JEPA 2 outperforms the previous state-of-the-art model PlausiVL by a significant margin, even with its 300 million parameters compared to the 8 billion parameters used in PlausiVL. In particular, V-JEPA 2 ViT-g384 demonstrates a \(+12.1\) points improvement over PlausiVL on action recall-at-5, corresponding to a \(44\%\) relative improvement.
在图 11 中,我们可视化了 V-JEPA 2 在 EK100 验证集三个样本上的预测结果,其中两个样本模型预测成功,一个样本预测失败。对于两个成功的示例,V-JEPA 2 不仅以最高置信度正确检索出动作,而且基于给定上下文提出了连贯的 top 2 到 top 5 动作。例如,在第一行中,正确动作是“清洗水槽”,但考虑到存在水龙头和墙壁,“打开水”或“清洁墙壁”也都是合理的动作。模型还预测了“冲洗海绵”,这是当前正在执行的动作,可能假设该动作在 1 秒后仍在进行。对于失败案例,V-JEPA 2 仍然提出了连贯的动作,如“关门”和“放下调料包”,但未能准确识别物体的具体性质:“茶包”。
In Figure 11 we visualize V-JEPA 2 predictions on three samples from the EK100 validation set, two where the model is successful and one where the model fails. For both successful examples, V-JEPA 2 not only retrieves the correct action with top 1 confidence, but also proposes coherent top 2 to 5 actions, based on the given context. For example, in the top row, the correct action is "wash sink", but "turn on water" or "clean wall" would both have been valid actions given the presence of a tap and a wall. The model also predicts "rinse sponge", which is the current action being performed, probably assuming that this action could still be going on after 1 second. For the failure case, V-JEPA 2 still proposes coherent actions such as "close door" and "put down spices package", but misses the exact nature of the object: "tea package".
V-JEPA 2 和 EK100 基准存在若干局限性。首先,V-JEPA 2 并未完全解决 EK100 任务,存在模型将动词、名词或两者都预测错误的失败案例。我们在 D.2 节中研究了这些失败的分布。其次,我们这里专注于预测 1 秒前瞻时间的动作。当预测更长的时间跨度时,V-JEPA 2 的准确率会下降,参见 D.2 节。第三,EK100 基准仅限于厨房环境,且词汇表封闭且定义明确,我们不知道 V-JEPA 2 在其他环境中的泛化能力如何。这限制了在 EK100 上训练的模型的实用性和适用性。最后,EK100 中的动作是从固定类别集合中选取的,因此无法泛化到训练集中未出现的动作类别。
V-JEPA 2 and the EK100 benchmark have several limitations. First, V-JEPA 2 does not fully solve EK100; there are failure cases where the model either gets the verb, the noun, or both wrong. We study the distribution of these failures in Section D.2. Second, we focus here on predicting actions with a 1 second anticipation time. The accuracy of V-JEPA 2 degrades when predicting at longer time horizons, see Section D.2. Third, the EK100 benchmark is limited to kitchen environments, with a closed well-defined vocabulary, and we do not know how well V-JEPA 2 generalizes to other environments. This limits the utility and applicability of models trained on EK100. Lastly, actions in EK100 are chosen from a fixed set of categories, making it impossible to generalize to action categories not present in the training set.
在本节中,我们探讨 V-JEPA 2 在开放语言视频问答(VidQA)任务上的能力。为了赋予语言能力,我们训练了一个多模态大语言模型(MLLM),采用 V-JEPA 2 作为视觉编码器,并使用由 LLaVA 模型系列推广的非分词早期融合设置。在这类 MLLM 中,视觉编码器通过将视觉编码器的输出补丁嵌入投影到 LLM 的输入嵌入空间,从而与大语言模型对齐。然后,MLLM 以端到端方式或冻结视觉编码器的方式进行训练。用于 VidQA 的 MLLM 中大多数编码器通常是图像编码器,它们对视频输入逐帧独立应用。这类编码器的常见实例包括 CLIP、SigLIP 和 Perception Encoder,它们主要因其通过图像-文本对预训练获得的语义对齐而入选。据我们所知,我们的工作是首次使用在没有语言监督的情况下预训练的视频编码器来训练用于 VidQA 的 MLLM。
In this section, we explore V-JEPA 2's ability to perform open-language video question answering (VidQA). To enable language capabilities, we train a Multimodal Large Language Model (MLLM) using V-JEPA 2 as the visual encoder in the non-tokenized early fusion setup popularized by the LLaVA family of models. In this family of MLLMs, a visual encoder is aligned with a large language model by projecting the output patch embeddings of the vision encoder to the input embedding space of the LLM. The MLLM is then trained either end-to-end, or with a frozen vision encoder. The majority of the encoders used in MLLMs for VidQA are typically image encoders, which are applied independently per-frame for video inputs. Popular instances of such encoders are CLIP, SigLIP, and Perception Encoder, which are chosen primarily due to their semantic alignment with language, obtained by pretraining with image-caption pairs. To the best of our knowledge, our work is the first to use a video encoder that is pretrained without any language supervision, to train an MLLM for VidQA.
MLLM 在下游任务上的性能也高度依赖于对齐数据。在这些实验中,我们使用了包含 8850 万图像-文本和视频-文本对的数据集,类似于训练 PerceptionLM 所用的数据集。为了展示 V-JEPA 2 编码器的有效性,我们首先在第 7.2 节中,在受控数据设置下,使用 1800 万样本的子集,将 V-JEPA 2 与其他最先进的视觉编码器进行比较。然后,在相同的受控设置下,我们在第 7.3 节中表明,扩展视觉编码器和输入分辨率大小都能持续提高 VidQA 性能。最后,我们扩展对齐数据,使用全部 8850 万样本,在第 7.4 节中测试 V-JEPA 2 语言对齐的极限。我们的结果表明,在受控数据设置下,V-JEPA 2 在开放式 VidQA 任务上与其他视觉编码器相比取得了具有竞争力的性能。在扩展对齐数据后,V-JEPA 2 在多个 VidQA 基准上达到了最先进的性能。
MLLM performance on downstream tasks is also highly dependent on the alignment data. In these experiments we use a dataset of 88.5 million image- and video-text pairs, similar to what was used to train PerceptionLM. To demonstrate the effectiveness of the V-JEPA 2 encoder, first we compare V-JEPA 2 with other state-of-the-art vision encoders in a controlled data setup in Section 7.2, using a subset of 18 million samples. Then, in the same controlled setup, we show that scaling the vision encoder and input resolution size both consistently improve VidQA performance in Section 7.3. Finally, we scale the alignment data, using the full 88.5 million samples to test the limits of language alignment with V-JEPA 2 in Section 7.4. Our results demonstrate that in a controlled data setup, V-JEPA 2 obtains competitive performance on open-ended VidQA tasks compared to other vision encoders. Upon scaling the alignment data, V-JEPA 2 achieves state-of-the-art performance on several VidQA benchmarks.
我们在 PerceptionTest 上进行评估,该测试评估模型在不同技能上的表现,如记忆、抽象、物理和语义。此外,我们在 MVP 数据集上评估物理世界理解能力,该数据集采用最小视频对评估框架以减轻文本和外观偏差。我们还使用 TempCompass、TemporalBench 和 TOMATO 来研究模型的时间理解和记忆能力。最后,我们使用 MVBench(其偏向于单帧外观特征)和 TVBench(文献中提出的替代方案,用于通用和时间理解,以减轻这些偏差)报告通用理解能力的结果。
We evaluate on PerceptionTest, which assesses model performance across different skills such as memory, abstraction, physics, and semantics. Additionally, we evaluate on the MVP dataset for physical world understanding, which utilizes a minimal-video pair evaluation framework to mitigate text and appearance biases. We also evaluate on TempCompass, TemporalBench, and TOMATO to investigate temporal understanding and memory capabilities of models. Finally, we report results on general understanding ability using MVBench, which has a bias towards single-frame appearance features, and TVBench, which is proposed in the literature as an alternative for general and temporal understanding, mitigating those biases.
为了评估 V-JEPA 2 表示在视觉问答任务上的表现,我们使用 LLaVA 框架中的视觉指令微调流程将 V-JEPA 2 与 LLM 对齐。该过程涉及使用可学习的投影模块(通常是 MLP)将视觉编码器输出(或视觉标记)转换为 LLM 输入。我们按照以下渐进式三阶段流程训练 MLLM:阶段 1,仅使用图像描述数据训练投影器;阶段 2,在大规模图像问答上训练完整模型;阶段 3,进一步在大规模视频描述和问答上训练模型。通过这种分阶段训练方法,LLM 逐步提高对视觉标记的理解。视觉编码器可以冻结,也可以与 MLLM 的其余部分一起微调。我们探索了这两种设置,因为冻结视觉编码器可以更清晰地反映视觉特征的质量,而微调视觉编码器则能获得更好的整体性能。视觉指令训练的更多细节见附录 E。
To evaluate the V-JEPA 2 representations on visual-question answering tasks, we align V-JEPA 2 with an LLM using the visual instruction tuning procedure from the LLaVA framework. This process involves converting the visual encoder outputs (or visual tokens) into LLM inputs using a learnable projector module, which is typically an MLP. We train MLLMs through a progressive three-stage process following: Stage 1, where we train the projector solely on image captioning data; Stage 2, where we train the full model on large-scale image question answering; and Stage 3, where we further train the model on large-scale video captioning and question answering. Through this staged training approach, the LLM incrementally improves its understanding of visual tokens. The vision encoder can either be frozen or finetuned along with the rest of the MLLM. We explore both settings, as freezing the vision encoder gives a cleaner signal about the quality of the visual features, while finetuning the vision encoder yields better overall performance. Further details of the visual instruction training are described in Appendix E.
为了隔离视觉编码器对多模态大语言模型(MLLM)性能的贡献,并与 V-JEPA 2 进行比较,我们引入了一个受控设置:使用相同的 LLM 主干和训练设置,用不同的最先进编码器训练单独的 MLLM。在此受控设置中,我们使用 Qwen2-7B-Instruct 并冻结视觉编码器。我们使用了 1800 万个图像和视频文本对齐样本。我们首先将预训练分辨率为 512 × 512 的 V-JEPA 2 与 DINOv2、SigLIP-2 和感知编码器(Perception Encoder)进行比较。
To isolate the contribution of vision encoders to MLLM performance and compare with V-JEPA 2, we introduce a controlled setup: we train individual MLLMs with different state-of-the-art encoders using the same LLM backbone and training setup. In this controlled setup, we use Qwen2-7B-Instruct and freeze the vision encoder. We use 18 million image and video-text aligned samples. We first compare V-JEPA 2, pretrained at resolution 512 × 512, with DINOv2, SigLIP-2, and Perception Encoder.
我们观察到,在冻结设置中,V-JEPA 2 表现出具有竞争力的性能,在所有测试基准(表 6)上均优于 DINOv2、SigLIP 和感知编码器(PE),但在 PerceptionTest 上,V-JEPA 2 略逊于 SigLIP 和 PE。改进在 MVP、TemporalBench 和 TVBench 上尤为明显——这些基准主要关注时间理解。此外,由于我们仅更换了视觉编码器,我们提供了证据表明,在没有语言监督的情况下训练的视频编码器可以胜过在有语言监督下训练的编码器,这与传统观念相反。结果还表明,在 VidQA 中使用视频编码器而非图像编码器可以提高时空理解能力,这凸显了开发更好的视频编码器的必要性。
We observe that V-JEPA 2 exhibits competitive performance in the frozen setup, outperforming DINOv2, SigLIP, and Perception Encoder (PE) in all of the tested benchmarks (Table 6) except PerceptionTest where V-JEPA 2 slightly underperforms SigLIP and PE. The improvement is especially noticeable on MVP, TemporalBench, and TVBench — benchmarks that are primarily focused on temporal understanding. Additionally, since we only change the vision encoder, we provide evidence that a video encoder trained without language supervision can outperform encoders trained with language supervision, in contrast to conventional wisdom. The results also indicate that using a video encoder instead of an image encoder for VidQA improves spatiotemporal understanding, highlighting the need to develop better video encoders.
先前的研究表明,缩放视觉编码器和输入分辨率能显著提升自监督图像编码器的 VQA 性能。因此,我们将 V-JEPA 2 从 300M 参数扩展到 1B 参数,并将输入分辨率从 256 像素提高到 512 像素,结果如表 7 所示。当固定输入分辨率为 256 像素,将视觉编码器容量从 300M 增加到 1B 参数时,我们在 PerceptionTest 上提升了 0.9 个百分点,在 TVBench 上提升了 3.3 个百分点,在 MVBench 上提升了 1.2 个百分点。此外,将输入分辨率提高到 512 像素在所有下游任务上带来了进一步改进,例如在 PerceptionTest 上提升了 2.2 个百分点,在 TemporalBench 上提升了 4.0 个百分点,在 TVBench 上提升了 3.3 个百分点。这些结果表明,进一步缩放视觉编码器和输入分辨率是提升 VidQA 性能的一个有前景的方向。
Prior work suggests that scaling the vision encoder and input resolution significantly improves VQA performance for self-supervised image encoders. Thus, we scale V-JEPA 2 from 300M to 1B parameters and the input resolution from 256 to 512 pixels, and show the results in Table 7. When increasing vision encoder capacity from 300M to 1B parameters for a fixed input resolution of 256 pixels, we observe improvements of 0.9 points on PerceptionTest, 3.3 points on TVBench, and 1.2 points on MVBench. Additionally, increasing the input resolution to 512 pixels yields further improvements across all downstream tasks, such as an improvement of 2.2 points on PerceptionTest, 4.0 points on TemporalBench, and 3.3 points on TVBench. These results suggest that further scaling the vision encoder and input resolution is a promising direction for improving VidQA performance.
在受控设置下对 V-JEPA 2 训练多模态大语言模型(MLLM)的能力有了更深入的理解之后,我们研究了增加对齐数据集规模对提升 VidQA 最先进水平的影响。下游任务性能的阶跃式提升通常通过增加训练数据的规模来实现,如 [引用] 所示。为此,我们将 MLLM 训练数据的规模从 1800 万增加到完整的 8850 万(4.7 倍)。虽然提高模型分辨率有助于提升下游性能,但也带来了在 LLM 输入中容纳大量视觉 token 的挑战。因此,我们选择了 V-JEPA 2 ViT-g384,每帧产生 288 个视觉 token。我们遵循与训练 V-JEPA 2 ViT-g384 相同的方案,使用 Llama 3.1 作为骨干网络。为简化训练过程,我们使用不带池化的 MLP 投影器。关于扩展训练设置的详细信息见附录 E。
After developing a better understanding of the capabilities of V-JEPA 2 for training an MLLM in the controlled setup, we study the effect of increasing alignment dataset size to improve the state-of-the-art of VidQA. Step changes on downstream task performance are often achieved by increasing the scale of the training data, as observed by [citation]. To that end, we increase the scale of MLLM training data from 18 million to the full 88.5 million (4.7×). While increasing the model resolution helps in downstream performance, it comes with the challenge of accommodating a large number of visual tokens in the LLM input. We therefore choose V-JEPA 2 ViT-g384, leading to 288 visual tokens per frame. We follow the same recipe as to train V-JEPA 2 ViT-g384, using Llama 3.1 as the backbone. To simplify the training process, we use an MLP projector without pooling. Details on the scaling training setup are described in Appendix E.
数据的统一扩展一致地提升了下游基准性能,在多个基准上取得了最先进的结果(表 8)——包括 PerceptionTest、MVP、TempCompass、TemporalBench 和 TOMATO。与当前最先进的 PerceptionLM 8B 相比,我们在 PerceptionTest 测试集上的准确率提升了 1.3 个百分点,在 MVP 上的配对准确率提升了 4.8 个百分点,在 TempCompass 上的准确率提升了 4.2 个百分点,在 TemporalBench 的短问答片段上的多二元准确率提升了 8.4 个百分点,在 TOMATO 上的准确率提升了 7.1 个百分点。V-JEPA 2 在 TVBench 和 MVBench 上未超过 PerceptionLM,但仍显著优于其他相关基线(InternVL 2.5、Qwen2VL 和 Qwen2.5VL)。这些结果强调了扩展训练数据规模对于视觉-语言对齐的必要性,并提供了证据表明,像 V-JEPA 2 这样在没有语言监督的情况下预训练的编码器,在足够的数据规模下也能达到最先进的性能。
Scaling the data uniformly improves the downstream benchmark performance, resulting in state-of-the-art results (Table 8) on multiple benchmarks — PerceptionTest, MVP, TempCompass, TemporalBench and TOMATO. Compared to the current state-of-the-art PerceptionLM 8B, we observe an increase of 1.3 points on accuracy for PerceptionTest test set, 4.8 points on paired accuracy for MVP, 4.2 points on accuracy for TempCompass, 8.4 points on Multi-binary accuracy for short-QA segment for TemporalBench and 7.1 points on accuracy for TOMATO. V-JEPA 2 does not outperform PerceptionLM on TVBench and MVBench, however it still significantly outperforms other related baselines (InternVL 2.5, Qwen2VL and Qwen2.5VL). These results underscore the need to scale training data for vision-language alignment and provide evidence that an encoder pretrained without language supervision, such as V-JEPA 2, can achieve state-of-the-art results with sufficient scale.
早在[某些]工作中,AI 研究者就致力于构建使用内部世界模型的智能体——既建模世界动态,也映射静态环境——以实现高效的规划和控制。以往研究已在模拟任务以及真实世界的移动和操作任务中探索了世界模型。世界模型方法要么直接在像素空间学习预测模型,要么在学习的表示空间中学习,要么利用更结构化的表示空间(如关键点表示)。以往在机器人任务中展示真实世界性能的方法都训练了特定任务的世界模型,并且依赖于机器人部署环境中的交互数据。评估的重点是在所探索的任务空间内展示世界建模方法的性能,而不是泛化到新环境或未见过的物体。在本工作中,我们训练了一个任务无关的世界模型,并展示了其对新环境和物体的泛化能力。
As early as the work of [and], AI researchers have sought to build agents that use internal models of the world — modeling both dynamics of the world, as well as mapping the static environment — to enable efficient planning and control. Previous work has investigated world models in simulated tasks, as well as real-world locomotion and manipulation tasks. World model approaches either learn predictive models directly in pixel-space, in a learned representation space, or utilizing more structured representation spaces such as keypoint representations. Previous approaches that have demonstrated real-world performance on robotics tasks have trained task-specific world models, and they rely on interaction data from the environment in which the robot is deployed. Evaluation is focused on demonstrating performance of world modeling approaches within the explored task space, instead of generalization to new environments or unseen objects. In this work we train a task-agnostic world model, and demonstrate generalization to new environments and objects.
一些近期工作利用互联网规模的视频和交互数据来训练面向自主机器人的通用(任务无关)动作条件视频生成模型。然而,迄今为止这些方法仅展示了在给定机器人动作时生成视觉上看似有效的计划的能力,但尚未展示使用这些模型实际控制机器人的能力。
Some recent works leverage both internet-scale video and interaction data towards training general purpose (task-agnostic) action-conditioned video generation models for autonomous robots. However, thus far these approaches only demonstrate the ability to generate visually valid-looking plans given actions of the robot, but they have not demonstrated the ability to use those models to actually control the robot.
其他工作探索了将生成建模整合到策略学习中。与这一系列工作不同,我们的目标是通过模型预测控制利用世界模型,而不是策略学习,以避免需要专家轨迹的模仿学习阶段。这两种方法是正交的,未来可以结合。与我们的工作最接近的是,[他们]表明可以分阶段或端到端地学习世界模型,并用它零样本解决规划任务。虽然那些先前的工作侧重于小规模规划评估,但我们表明类似的原理可以扩展并用于解决真实世界的机器人任务。
Other works have explored the integration of generative modeling into policy learning. Differently from this line of work, our goal is to leverage a world model through model-predictive control instead of policy learning to avoid the imitation learning phase that requires expert trajectories. Both approaches are orthogonal and could be combined in future works. Closest to our work, [they] show that you can learn a world model stage-wise or end-to-end and use it to solve planning tasks zero-shot. While those previous works focus on small-scale planning evaluation, we show that similar principles can be scaled and used to solve real-world robotic tasks.
近期真实世界机器人控制中的模仿学习方法在学习具有越来越强泛化能力的策略方面取得了显著进展。这是通过利用在互联网规模的视频和文本数据上预训练的视频语言模型实现的,这些模型随后通过从专家示范中进行行为克隆来微调(或适应)以预测动作。尽管这些方法展示了有前景的泛化结果,但尚不清楚它们是否能学会预测训练数据中未示范的行为,因为它们缺乏显式的世界预测模型,并且不利用推理时的计算进行规划。它们需要高质量的大规模遥操作数据,并且只能利用成功的轨迹。相比之下,我们专注于利用任何交互数据,无论其来自与环境的成功还是失败的交互。
Recent imitation learning approaches in real-world robotic control have made significant progress towards learning policies that show increasingly good generalization capabilities. This is achieved by leveraging video-language models that have been pre-trained on internet-scale video and text data, which are then fine-tuned (or adapted) to also predict actions by using behavior cloning from expert demonstrations. Although these approaches show promising generalization results, it is unclear whether they will be able to learn to predict behaviors that were not demonstrated in the training data since they lack an explicit predictive model of the world and do not leverage inference-time computation for planning. They require high-quality large-scale teleoperation data, and can only utilize successful trajectories. In contrast, we focus on leveraging any interaction data whether it comes from a successful or failed interaction with the environment.
计算机视觉中的视频基础模型已经表明,由图像和/或视频组成的大规模观测数据集可用于学习通用的视觉编码器,这些编码器通过自监督学习方法(来自图像、视频、弱语言监督或其组合)在广泛的下游任务中表现良好。然而,以往的工作往往侧重于在与大语言模型对齐后使用基于探针的评估或视觉问答任务来理解评估。虽然这些任务推动了进展,但视觉系统的一个重要目标仍然是使智能体能够与物理世界交互。除了视觉理解任务的结果外,我们还研究了大规模视频自监督学习如何能够以零样本方式解决新环境中的规划任务。
Video foundation models in computer vision have shown that large-scale observation datasets comprised of images and/or videos can be leveraged to learn generalist vision encoders that perform well along a wide range of downstream tasks using self-supervised learning approaches from images, videos, with weak language supervision, or a combination thereof. Previous works, however, tend to focus on understanding evaluation using probe-based evaluation or visual question answering tasks after aligning with a large-language model. While such tasks have served to drive progress, it remains an important goal of a visual system to enable an agent to interact with the physical world. Beyond results on visual understanding tasks, we investigate how large-scale self-supervised learning from video can enable solving planning tasks in new environments in a zero-shot manner.
本研究表明,联合嵌入预测架构通过从网络规模数据和少量机器人交互数据中进行自监督学习,能够产生一个能够理解、预测和规划物理世界的世界模型。V-JEPA 2 在需要运动理解和人类动作预测的动作分类任务上取得了最先进的性能。当与大型语言模型对齐时,V-JEPA 2 在视频问答任务上也优于之前的视觉编码器。此外,使用 V-JEPA 2 的表示对动作条件世界模型 V-JEPA 2-AC 进行后训练,使得真实世界机器人能够成功完成零样本的抓取操作任务,例如拾取和放置。这些发现表明 V-JEPA 2 是朝着开发能够有效感知和作用于其环境的先进人工智能系统迈出的一步。
This study demonstrates how joint-embedding predictive architectures, learning in a self-supervised manner from web-scale data and a small amount of robot interaction data, can yield a world model capable of understanding, predicting, and planning in the physical world. V-JEPA 2 achieves state-of-the-art performances on action classification requiring motion understanding and human action anticipation. V-JEPA 2 also outperforms previous vision encoders on video question-answering tasks when aligned with a large language model. Additionally, post-training an action-conditioned world model, V-JEPA 2-AC, using V-JEPA 2’s representation, enables successful zero-shot prehensile manipulation tasks, such as Pick-and-Place, with real-world robots. These findings indicate V-JEPA 2 is a step towards developing advanced AI systems that can effectively perceive and act in their environment.
未来工作有几个重要方向来解决 V-JEPA 2 的局限性。首先,在这项工作中,我们专注于需要预测未来约 16 秒的任务。这使得从单个目标图像进行更简单的操作任务(如抓取和带物体到达)的规划成为可能。然而,要将其扩展到更长时域的任务(如拾取和放置)甚至更复杂的任务,而不需要子目标,将需要建模方面的进一步创新。开发能够在多个空间和时间尺度、不同抽象层次上进行预测的分层模型方法是一个有前景的方向。
There are several important avenues for future work to address limitations of V-JEPA 2. First, in this work we have focused on tasks requiring predictions up to roughly 16 seconds into the future. This enables planning for simpler manipulation tasks, like grasp and reach-with-object, from a single goal image. However, to extend this to longer-horizon tasks such as pick-and-place or even more complex tasks, without requiring sub-goals will require further innovations in modeling. Developing approaches for hierarchical models capable of making predictions across multiple spatial and temporal scales, at different levels of abstraction, is a promising direction.
其次,如第 4 节所述,V-JEPA 2-AC 目前依赖于以图像目标指定的任务。虽然这对某些任务可能是自然的,但在其他情况下,基于语言的目标指定可能更可取。扩展 V-JEPA 2-AC 以接受基于语言的目标,例如,通过一个能够将基于语言的目标嵌入到 V-JEPA 2-AC 表示空间中的模型,是未来工作的另一个重要方向。第 7 节中描述的结果,即将 V-JEPA 2 与语言模型对齐,可以作为起点。
Second, as mentioned in Section 4, V-JEPA 2-AC currently relies upon tasks specified as image goals. Although this may be natural for some tasks, there are other situations where language-based goal specification may be preferable. Extending the V-JEPA 2-AC to accept language-based goals, e.g., by having a model that can embed language-based goals into the V-JEPA 2-AC representation space, is another important direction for future work. The results described in Section 7, aligning V-JEPA 2 with a language model, may serve as a starting point.
最后,在这项工作中,我们将 V-JEPA 2 模型扩展到适度的 1B 参数。第 2 节的结果表明,在扩展到这一水平时,性能持续提升。先前的工作已经研究了将视觉编码器扩展到多达 20B 参数。需要在这一方向上开展更多工作,以开发可扩展的预训练方法,从而随着规模扩大实现持续的性能提升。
Finally, in this work we scaled V-JEPA 2 models up to a modest 1B parameters. The results in Section 2 demonstrated consistent performance improvements while scaling to this level. Previous work has investigated scaling vision encoders to as large as 20B parameters. Additional work is needed in this direction to develop scalable pre-training recipes that lead to sustained performance improvements with scale.