3D-aware image synthesis encompasses a variety of tasks, such as scene generation and novel view synthesis from images. Despite numerous task-specific methods, developing a comprehensive model remains challenging. In this paper, we present SSDNeRF, a unified approach that employs an expressive diffusion model to learn a generalizable prior of neural radiance fields (NeRF) from multi-view images of diverse objects. Previous studies have used two-stage approaches that rely on pretrained NeRFs as real data to train diffusion models. In contrast, we propose a new single-stage training paradigm with an end-to-end objective that jointly optimizes a NeRF auto-decoder and a latent diffusion model, enabling simultaneous 3D reconstruction and prior learning, even from sparsely available views. At test time, we can directly sample the diffusion prior for unconditional generation, or combine it with arbitrary observations of unseen objects for NeRF reconstruction. SSDNeRF demonstrates robust results comparable to or better than leading task-specific methods in unconditional generation and single/sparse-view 3D reconstruction.

本文提出了一种称为SSDNeRF的新方法，它使用表达能力强的Diffusion Model从多视图图像中学习神经辐射场（NeRF）的可推广先验，实现3D重建和先验学习的同时, 证明了该方法在无条件生成和单/稀疏视图3D重建等任务上具有与任务特定方法媲美或优于其的鲁棒性结果。

单级扩散NeRF：一种统一的三维生成和重建方法