SV3D:基于潜在视频扩散的单图像新颖多视图合成与 3D 生成

SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion

罗宾·罗姆巴赫 Robin Rombach · Stability AI · 2024-03-18 · arXiv:2403.12008 ↗ · 被引 402

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了 Stable Video 3D(SV3D)——一种潜在视频扩散模型,用于从单张图像生成围绕 3D 物体的高分辨率、多视角轨道视频。最近的 3D 生成工作提出了将 2D 生成模型适配到新视角合成(NVS)和 3D 优化的技术。然而,这些方法由于视角有限或 NVS 不一致而存在若干缺点,从而影响 3D 物体生成的性能。在这项工作中,我们提出了 SV3D,它将图像到视频扩散模型适配到新颖多视图合成和 3D 生成,从而利用了视频模型的泛化能力和多视图一致性,同时进一步为 NVS 添加了显式相机控制。我们还提出了改进的 3D 优化技术,以使用 SV3D 及其 NVS 输出进行图像到 3D 生成。在多个数据集上的大量实验结果(包括 2D 和 3D 指标以及用户研究)表明,与先前工作相比,SV3D 在 NVS 和 3D 重建方面达到了最先进的性能。

We present Stable Video 3D (SV3D) -- a latent video diffusion model for high-resolution, image-to-multi-view generation of orbital videos around a 3D object. Recent work on 3D generation propose techniques to adapt 2D generative models for novel view synthesis (NVS) and 3D optimization. However, these methods have several disadvantages due to either limited views or inconsistent NVS, thereby affecting the performance of 3D object generation. In this work, we propose SV3D that adapts image-to-video diffusion model for novel multi-view synthesis and 3D generation, thereby leveraging the generalization and multi-view consistency of the video models, while further adding explicit camera control for NVS. We also propose improved 3D optimization techniques to use SV3D and its NVS outputs for image-to-3D generation. Extensive experimental results on multiple datasets with 2D and 3D metrics as well as user study demonstrate SV3D's state-of-the-art performance on NVS as well as 3D reconstruction compared to prior works.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 10)

阅读逐段中英对照全文 →