Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Why does Deep Learning Improve Visual SLAM?

About

Visual SLAM is a well-established technology utilized in a wide range of real-world applications. However, its performance still degrades under challenging visual conditions, such as low texture, severe motion blur, and poor illumination. Systems based on deep learning outperform classical geometry-based ones and achieve state-of-the-art results by combining learned 2D data association and uncertainty with differentiable geometric optimization in recurrent architectures. Still, it remains unclear exactly which components are fundamentally responsible for this success. In this paper, we ask: Is the superior performance of deep learning-based systems driven primarily by learned 2D data association, the combination of learned 2D data association and uncertainty, or the recurrent architecture itself? We investigate this question empirically by conducting a controlled study. Our findings reveal that the success of DL-based V-SLAM systems hinges on learned 2D data association and uncertainty rather than their recurrent architecture, underscoring the necessity of learning-based paradigms for the design of these components. Upon acceptance, the code will be released as open source.

Giovanni Cioffi, Davide Scaramuzza• 2026

Related benchmarks

TaskDatasetResultRank
Visual SLAMTartanAir
ATE Translation (m)0.04
104
Trajectory EstimationUZH-FPV (All Sequences)
ATE Translation (m)0.05
64
Visual SLAMUZH-FPV Ind. Forw. 3
Average Time per Frame (sec)0.051
7
Visual SLAMTartanAir ME007
Average Time per Frame (s)0.05
7
Showing 4 of 4 rows

Other info

Follow for update