BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots
About
Scale-consistent ego-motion estimation is fundamental for autonomous ground robots. Bird's-Eye-View (BEV) representation naturally addresses the scale drift problem of monocular visual odometry (MVO) by providing a metric-scaled planar workspace, enabling the simplification of 6-DoF ego-motion to a more robust 3-DoF model. However, existing BEV-based methods suffer from two key limitations: sparse supervision signals from pose-only training, and information loss during perspective-to-BEV projection. We present BEV-ODOM2, an enhanced framework that addresses both limitations without requiring additional annotations. Our approach introduces (1) dense BEV optical flow supervision constructed directly from 3-DoF pose ground truth for pixel-level guidance, and (2) Perspective View (PV)-BEV fusion that computes correlation volumes before projection to preserve 6-DoF motion cues. An enhanced rotation sampling strategy further balances diverse motion patterns during training. We evaluate on four datasets with varied spatial scales: KITTI, Oxford, NCLT, and our newly collected ZJH-VO benchmark. BEV-ODOM2 achieves a 40\% RTE improvement over prior BEV-based methods, with real-time inference on an NVIDIA Jetson AGX Orin confirming edge deployment feasibility. The source code and the ZJH-VO dataset are publicly released to facilitate future research.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Visual Odometry | KITTI Seq. 10 | Translational Error (%)3.46 | 35 | |
| Visual Odometry | Oxford | Trajectory Error (11-12)0.91 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-04) | ATE (m)1.55 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-00) | ATE (m)0.57 | 18 | |
| Visual Odometry | ZJH-VO Average multi-scale | ATE (m)2.04 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-09) | ATE (m)1.35 | 18 | |
| Visual Odometry | ZJH-VO multi-scale (F-01) | ATE (m)4.67 | 18 | |
| Visual Odometry | NCLT | Sequence 03-17 Error2.2 | 16 | |
| Visual Odometry | NCLT | FPS69.52 | 11 | |
| Lateral Dead-Reckoning | KITTI (Seq10) | Median Lateral Error (m)0.89 | 3 |