Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding

About

3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by generalization and recent zero-shot methods typically rely on a predefined Object Lookup Table (OLT) to query Visual Language Models (VLMs) for reasoning about object locations via a single step grounding, which limits the applications in scenarios with undefined targets and complex queries. To address these problems, we present OpenGround, a novel zero-shot framework for open-world 3D visual grounding that remains compatible with recent zero-shot methods. OpenGround integrates Task-Chain Planning to decompose a query into a plan of context-to-target sub-goals for progressive grounding, and Context-Guided Perception to perceive novel objects online under context guidance from the task chain. We also propose a new dataset named OpenTarget, which contains over 7000 object-description pairs to mimic open-world evaluation. Extensive experiments demonstrate that OpenGround achieves competitive performance on Nr3D, state-of-the-art on ScanRefer, and delivers a substantial 17.6\% improvement on OpenTarget. Project Page at https://why-102.github.io/openground.io/.

Wenyuan Huang, Zhenyu Zhang, Zhao Wang, Zhou Wei, Ting Huang, Fang Zhao, Jian Yang• 2025

Related benchmarks

TaskDatasetResultRank
3D Visual GroundingScanRefer
Acc@0.2557.9
172
3D Visual GroundingNr3D
Overall Success Rate61.7
109
3D Visual GroundingScanRefer Overall
Acc @ 0.553.1
55
3D Visual GroundingScanRefer Unique
Acc@0.25 (IoU=0.25)77.8
41
3D Visual GroundingOpenTarget randomly selected 300 samples
Accuracy @ IoU=0.2546.2
6
Showing 5 of 5 rows

Other info

Follow for update