BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation
About
Monocular visual odometry (MVO) is vital in autonomous navigation and robotics, providing a cost-effective and flexible motion tracking solution, but the inherent scale ambiguity in monocular setups often leads to cumulative errors over time. In this paper, we present BEV-ODOM, a novel MVO framework leveraging the Bird's Eye View (BEV) Representation to address scale drift. Unlike existing approaches, BEV-ODOM integrates a depth-based perspective-view (PV) to BEV encoder, a correlation feature extraction neck, and a CNN-MLP-based decoder, enabling it to estimate motion across three degrees of freedom without the need for depth supervision or complex optimization techniques. Our framework reduces scale drift in long-term sequences and achieves accurate motion estimation across various datasets, including NCLT, Oxford, and KITTI. The results indicate that BEV-ODOM outperforms current MVO methods, demonstrating reduced scale drift and higher accuracy.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Visual Odometry | KITTI Seq. 10 | Translational Error (%)3.61 | 35 | |
| Visual Odometry | ZJH-VO multi-scale (F-09) | ATE (m)1.12 | 18 | |
| Visual Odometry | Oxford | Trajectory Error (11-12)1.38 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-04) | ATE (m)1.91 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-00) | ATE (m)0.77 | 18 | |
| Visual Odometry | ZJH-VO Average multi-scale | ATE (m)3.68 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-01) | ATE (m)10.94 | 18 | |
| Visual Odometry | NCLT | Sequence 03-17 Error3.73 | 16 | |
| Visual Odometry | NCLT | FPS92.84 | 11 | |
| Lateral Dead-Reckoning | KITTI (Seq09) | Median Lateral Error (m)0.83 | 3 |