SurroundNEXO: Ego-Centric Metric Bridging for Spatially Consistent Geometry in Autonomous Driving
About
Modern autonomous driving depends on accurate metric 3D understanding for perception, reconstruction, and planning, which in turn requires reliable multi-camera depth prediction. However, the outward-facing nature of vehicle-mounted surround-view camera rigs inherently limits visual overlap across views, challenging the correspondence-based assumptions that underpin conventional multi-view geometry. To bridge this gap, we present SurroundNEXO, named after the Spanish word nexo for a geometric link, a low-overlap multi-camera metric depth framework that grounds cross-view reasoning in ego-centric geometry rather than dense visual correspondences. Instead of directly enforcing early global fusion, SurroundNEXO first assigns image tokens globally comparable ego-frame viewing directions through Ego-Ray Positional Encoding, then uses sparse LiDAR measurements as metric anchors to propagate absolute scale cues, and finally expands feature interaction progressively from view-local modeling to decomposed spatio-temporal reasoning and global integration. This design enables metric-scale depth prediction with improved spatial consistency across weakly overlapping cameras. Across low-overlap autonomous driving benchmarks, including NuScenes, Waymo and DDAD, SurroundNEXO reduces single-view error by 33.2%, improves cross-view consistency by 10.5%, and enhances metric reconstruction quality by 25.6% compared with SOTA methods. It further remains robust under extremely sparse depth prompts and exhibits strong zero-shot generalization to unseen camera layouts.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| 3D Reconstruction | DDAD | Accuracy0.221 | 36 | |
| Monocular Depth Estimation | DDAD | Abs Rel Error0.077 | 33 | |
| 3D Reconstruction | nuScenes | Acc0.281 | 29 | |
| 3D Reconstruction | OpenScene | Accuracy0.145 | 22 | |
| Single-view metric depth estimation | Waymo | Abs Rel0.048 | 20 | |
| Single-view metric depth estimation | OpenScene | Absolute Relative Error (Abs Rel)0.084 | 20 | |
| Depth Estimation | nuScenes | Abs Rel0.079 | 18 | |
| 3D Reconstruction | KITTI | -- | 17 | |
| Monocular Depth Estimation | Argoverse | Absolute Relative Error (AbsRel)0.081 | 14 | |
| Cross-view depth consistency | nuScenes | Absolute Relative Error0.139 | 12 |