Difix3D-W: Distractor-Free Few-Shot 3D Gaussian Splatting in the Wild
About
We propose Difix3D-W, a 3D novel sparse-view synthesis framework for unconstrained real-world scenarios that contain distractors, occlusion, and appearance variation. Unlike existing methods that primarily perform novel-view synthesis from a sparse set of constrained images without transient elements or leverage unconstrained dense image collections in real-world scenarios, our method utilize sparse unconstrained images, showing high-quality 3D rendering results. To do this, we introduce reference-guided view refinement with a redesigned one-step diffusion model using a transient mask and a reference image to mitigate artifacts in rendered views, enhancing the 3D representation in the Gaussian field. Furthermore, we address sparse regions in the Gaussian field leveraging sparsity-aware Gaussian replication strategy to amplify Gaussians in the sparse regions and alleviate deficient camera viewpoint issues. Finally, we utilize LoRA and regularization to maintain 3D multi-view consistency. Extensive experiments demonstrate that our method consistently outperforms existing methods. This advancement paves the way for realizing real-world scenarios without labor-intensive data acquisition.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Sparse-view 3D reconstruction | Photo Tourism sparse-view | PSNR19.86 | 11 | |
| Sparse-view 3D reconstruction | LLFF sparse-view | PSNR23.53 | 11 | |
| Sparse-view 3D reconstruction | NeRF On-the-go 3-view | PSNR25.23 | 8 | |
| Sparse-view 3D reconstruction | NeRF On-the-go 6-view | PSNR24.87 | 8 | |
| Sparse-view 3D reconstruction | NeRF On-the-go 9-view | PSNR25.08 | 8 | |
| Sparse-view 3D reconstruction | NeRF On-the-go Average | PSNR25.06 | 8 | |
| 3D Reconstruction | NeRF On-the-go 9-view training setting (test) | PSNR (Mountain)23.03 | 7 | |
| 3D Reconstruction | NeRF On-the-go 3-view (train) | Mountain PSNR24.66 | 7 |