A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove that LeJEPA (alignment plus Gaussian regularization) linearly recovers the world's latent variables from nonlinear observations, a property known as linear identifiability, in a broad class of worlds where latents evolve under stationary, additive-noise transitions. Our main result is that among all such worlds, the Gaussian is the unique latent distribution for which this guarantee holds. The forward direction rests on a spectral decomposition in which each degree of nonlinearity is strictly penalized by alignment, making the linear map the optimum; the converse rules out every non-Gaussian alternative. We further prove an approximate identifiability result where the guarantee degrades gracefully, and show that linear, orthogonal identifiability enables optimal latent-space planning. We validate the theory with experiments ranging from 2D examples to 1024-dimensional latents, including distributional ablations and pixel-based robotic control. Our theory turns an empirically successful recipe into a mathematical guarantee, providing the foundation for building World Models that provably recover the structure of the world.
核心贡献 · Key contributions
证明 LeJEPA(对齐+高斯正则化)在高斯潜在变量世界中实现线性可识别性。 Proves that LeJEPA (alignment + Gaussian regularization) achieves linear identifiability for Gaussian latent worlds.
确立高斯分布是在平稳加性噪声转移下保证线性可识别性的唯一潜在分布。 Establishes Gaussian as the unique latent distribution guaranteeing linear identifiability under stationary additive-noise transitions.
推导出近似可识别性界限,该界限随对齐和白化误差优雅退化。 Derives an approximate identifiability bound that degrades gracefully with alignment and whitening errors.
表明线性正交可识别性使得旋转不变成本下的最优潜在空间规划成为可能。 Shows linear orthogonal identifiability enables optimal latent-space planning for rotation-invariant costs.
在 2D 到 1024 维潜在变量、分布消融和基于像素的机器人控制中验证理论。 Validates theory across 2D to 1024-dimensional latents, distributional ablations, and pixel-based robotic control.
为联合嵌入预测架构(JEPA)提供了首个可识别性结果。 Provides the first identifiability result for Joint-Embedding Predictive Architectures (JEPAs).