LeJEPA 何时能学习世界模型?

When Does LeJEPA Learn a World Model?

杨立昆 Yann LeCun · NYU · 2026-05-25 · arXiv:2605.26379 ↗ · 被引 3

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

一种打乱世界真实自由度的表示无法支持可靠的规划或组合泛化。我们证明,在潜在变量遵循平稳加性噪声演化的广泛世界类别中,LeJEPA(对齐加高斯正则化)能从非线性观测中线性恢复世界的潜在变量,这一性质称为线性可识别性。我们的主要结果是:在所有此类世界中,高斯分布是唯一能保证这一性质的潜在分布。正向方向依赖于谱分解,其中每个非线性度都受到对齐的严格惩罚,使得线性映射成为最优;反向方向排除了所有非高斯替代方案。我们进一步证明了近似可识别性结果,其中保证会优雅地退化,并表明线性正交可识别性能够实现最优潜在空间规划。我们通过从二维示例到 1024 维潜在变量的实验验证了该理论,包括分布消融和基于像素的机器人控制。我们的理论将经验上成功的配方转化为数学保证,为构建能够可靠恢复世界结构的世界模型奠定了基础。

A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove that LeJEPA (alignment plus Gaussian regularization) linearly recovers the world's latent variables from nonlinear observations, a property known as linear identifiability, in a broad class of worlds where latents evolve under stationary, additive-noise transitions. Our main result is that among all such worlds, the Gaussian is the unique latent distribution for which this guarantee holds. The forward direction rests on a spectral decomposition in which each degree of nonlinearity is strictly penalized by alignment, making the linear map the optimum; the converse rules out every non-Gaussian alternative. We further prove an approximate identifiability result where the guarantee degrades gracefully, and show that linear, orthogonal identifiability enables optimal latent-space planning. We validate the theory with experiments ranging from 2D examples to 1024-dimensional latents, including distributional ablations and pixel-based robotic control. Our theory turns an empirically successful recipe into a mathematical guarantee, providing the foundation for building World Models that provably recover the structure of the world.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 20)

阅读逐段中英对照全文 →