Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs

About

Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph unqueryable during exploration. We argue that this sequential coupling is unnecessary and propose an asynchronous architecture in which lightweight online mapping runs concurrently with heavyweight semantic refinement. A probabilistic voxel-based backbone maintains stable object identities incrementally, while background VLM agents progressively enrich the graph. This framework resolves duplicate object tracks through semantic loop closure, attaches fine-grained visual attributes and derives spatial relations between objects. A multi-target frame scheduler amortizes VLM cost by selecting a small set of informative frames that jointly cover multiple targets. The resulting scene graph is queryable during exploration and grows in semantic richness over time. Our method matches or outperforms existing open-vocabulary 3D scene graph methods on semantic segmentation (ScanNet, Replica) and surpasses the prior state-of-the-art across three visual grounding benchmarks (Sr3D+, Nr3D, ScanRefer) by 15.3 to 18.8 A@0.25. Project page: https://denizbickici.github.io/thinkgraphs/

Deniz Bickici, Michael Pabst, Shohei Mori, Dieter Schmalstieg• 2026

Related benchmarks

TaskDatasetResultRank
3D Visual GroundingScanRefer
Acc@0.2552.9
172
3D Visual GroundingNr3D--
109
3D Semantic SegmentationReplica
3D mIoU37
61
3D Semantic SegmentationScanNet
Mean Accuracy (mAcc)75
7
3D Visual GroundingSr3D+
A@0.148.6
6
Showing 5 of 5 rows

Other info

Follow for update