Wild3R: Feed-Forward 3D Gaussian Splatting from Unconstrained Sparse Photo Collection
About
Feed-forward 3D Gaussian Splatting (3DGS) removes the need for time-consuming per-scene optimization required by traditional 3DGS. However, existing feed-forward approaches struggle with real-world photo collections that include diverse lighting conditions and transient objects. In this paper, we present Wild3R, a feed-forward approach for unconstrained sparse photo collections. The main bottleneck is the lack of training data that provides multiple viewpoints, a variety of illuminations, and transient variations necessary for learning robust scene representations. To address this, we introduce the WildCity dataset, which comprises 200 scenes, 170 lighting conditions, and transient objects, resulting in 337,500 images in total. By leveraging the dataset, our model learns appearance consistency across viewpoints conditioned on reference views, while removing transient content. Extensive experiments demonstrate that our method outperforms existing feed-forward approaches and achieves results competitive with prior per-scene optimization-based methods.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| View Synthesis | NeRF-OSR europa (test) | PSNR13.27 | 10 | |
| Novel View Synthesis | Photo Tourism 4 Context Views | PSNR13.04 | 10 | |
| Novel View Synthesis | Photo Tourism 16 Context Views | PSNR15.87 | 10 | |
| View Synthesis | NeRF-OSR stjohann (test) | PSNR12.22 | 10 | |
| Novel View Synthesis | Photo Tourism 64 Context Views | PSNR16.29 | 10 | |
| View Synthesis | NeRF-OSR st (test) | PSNR12.67 | 10 | |
| View Synthesis | NeRF-OSR lwp (test) | PSNR10.68 | 10 |