模态强制:可扩展的空间生成方法

Modality Forcing for Scalable Spatial Generation

贾斯汀·约翰逊 Justin Johnson · World Labs · 2026-06-11 · arXiv:2606.13676 ↗ · 被引 0

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

文本到图像(T2I)模型包含丰富的空间先验。合成逼真、杂乱的场景需要理解几何,包括透视和相对尺度。先前的工作将 T2I 模型适配以利用这种先验进行深度预测,但它们需要密集的深度数据并涉及复杂的方案。我们提出模态强制,一种简单、可扩展的后训练方案,使用在稀疏深度数据上训练的单个 DiT 进行联合图像-深度生成。模态强制通过为每种模态分配独立的噪声水平,支持图像和深度的条件生成和联合生成,且排列顺序任意。每种模态的解码器使我们能够在稀疏的真实世界深度上训练,并实现强大的、可泛化的深度预测。我们进一步表明,模态强制继承了 T2I 预训练的可扩展性:通过从头开始训练一组 T2I 模型(370M 到 3.3B 参数),我们发现,在更多图像数据上训练的更大模型能产生更准确的深度。我们最强的模型与最先进的单目深度估计器竞争,并将现有联合图像-深度生成模型的 AbsRel 降低了 57%。这些结果提供了强有力的证据,表明图像生成是空间感知的可扩展预训练目标。https://modality-forcing.github.io/

Text-to-image (T2I) models contain rich spatial priors. Synthesizing photorealistic, cluttered scenes requires an understanding of geometry, including perspective and relative scale. Prior works adapt T2I models to leverage this prior for depth prediction, but they require dense depth data and involve complex recipes. We propose Modality Forcing, a simple, scalable post-training recipe for joint image-depth generation using a single DiT trained on sparse depth data. Modality Forcing enables conditional and joint generation of image and depth in any permutation by assigning separate noise levels per modality. Per-modality decoders let us train on sparse, real-world depth and achieve strong, generalizable depth prediction. We further show that Modality Forcing inherits the scalability of T2I pre-training: by training a set of T2I models from scratch (370M to 3.3B parameters), we find that larger models trained on more image data produce more accurate depth. Our strongest model is competitive with state-of-the-art monocular depth estimators and reduces AbsRel by 57% relative to existing joint image-depth generative models. These results provide strong evidence that image generation is a scalable pre-training objective for spatial perception. https://modality-forcing.github.io/

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

阅读逐段中英对照全文 →