Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation

About

Recent advances in Iterative Vision-and-Language Navigation (IVLN) introduce a more meaningful and practical paradigm of VLN by maintaining the agent's memory across tours of scenes. Although the long-term memory aligns better with the persistent nature of the VLN task, it poses more challenges on how to utilize the highly unstructured navigation memory with extremely sparse supervision. Towards this end, we propose OVER-NAV, which aims to go over and beyond the current arts of IVLN techniques. In particular, we propose to incorporate LLMs and open-vocabulary detectors to distill key information and establish correspondence between multi-modal signals. Such a mechanism introduces reliable cross-modal supervision and enables on-the-fly generalization to unseen scenes without the need of extra annotation and re-training. To fully exploit the interpreted navigation data, we further introduce a structured representation, coded Omnigraph, to effectively integrate multi-modal information along the tour. Accompanied with a novel omnigraph fusion mechanism, OVER-NAV is able to extract the most relevant knowledge from omnigraph for a more accurate navigating action. In addition, OVER-NAV seamlessly supports both discrete and continuous environments under a unified framework. We demonstrate the superiority of OVER-NAV in extensive experiments.

Ganlong Zhao, Guanbin Li, Weikai Chen, Yizhou Yu• 2024

Related benchmarks

TaskDatasetResultRank
Vision-and-Language NavigationREVERIE (val unseen)
SPL22
129
Vision-and-Language NavigationREVERIE seen (val)
SR40
28
Iterative Vision-and-Language NavigationIR2R-CE (val seen)
TL9.5
15
Vision-and-Language NavigationGSA-R2R N-Scene (test)
SR16.7
14
Vision-and-Language NavigationGSA-R2R N-Basic (test)
TL11.4
10
Vision-and-Language NavigationGSA-R2R R-Basic (test)
Trajectory Length14.1
10
Vision-Language NavigationGSA-R2R Child instructions
SR20.9
10
Vision-Language NavigationGSA-R2R Keith instructions
SR20.5
10
Vision-Language NavigationGSA-R2R Moira instructions
SR19.5
10
Vision-Language NavigationGSA-R2R Rachel instructions
SR20.6
10
Showing 10 of 15 rows

Other info

Follow for update