NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
About
Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in long image sequences due to cumulative camera pose estimation drift, which propagates errors into geometric modeling and severely limits rendering fidelity. In this work, we revisit the long-sequence bottleneck and identify pose drift as the primary factor restricting reconstruction quality. Furthermore, while SfM-based pseudo ground-truth poses introduce sensor noise, purely rendering-based supervision often leads to optimization instability and local minima due to the entangled optimization of geometry and pose. To address the challenges, we propose a synergistic pose-free framework that explicitly couples geometry and appearance via a Raymap-Guided Coupling Module (RGC). Concretely, we anchor Gaussian centers to raymap-induced geometry and jointly optimize RGB reconstruction, raymap consistency, and camera regularization under a unified objective, yielding a bidirectional feedback loop: stronger geometry improves rendering, and appearance supervision in turn refines geometry and pose. To further stabilize learning across wide temporal ranges, we introduce a Dual-Frequency Viewpoint Scheduling strategy that combines easy-to-hard interval expansion with replay of short-interval pairs. Extensive experiments across in-domain and cross-domain datasets show consistent gains in both rendering and pose estimation, with notably improved robustness on long sequences. Ablation studies validate our central insight: explicitly designed geometry-appearance synergy is the key to scalable and drift-robust pose-free feed-forward 3D reconstruction. Project page: https://xiangyu1sun.github.io/NoDrift3R-project-page/
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Pose Estimation | RE10K | AUC @ 5°0.755 | 41 | |
| Novel View Synthesis | DL3DV 12 views | PSNR24.25 | 36 | |
| Novel View Synthesis | DL3DV 24 views | PSNR24.242 | 35 | |
| Novel View Synthesis | DL3DV 6 views | PSNR24.922 | 19 | |
| Novel View Synthesis | ScanNet++ 32v | PSNR17.569 | 14 | |
| Camera pose estimation | DL3DV 6 views | AUC@5°96.7 | 7 | |
| Novel View Synthesis | RE10k 6 input views (test) | PSNR25.736 | 6 | |
| Camera pose estimation | DL3DV 12 views | AUC@5°96.1 | 4 | |
| Camera pose estimation | DL3DV 24 views | AUC@5°94.9 | 4 | |
| Novel View Synthesis | ScanNet++ 64 views (long-sequence evaluation) | PSNR18.935 | 3 |