Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

SAT: 2D Semantics Assisted Training for 3D Visual Grounding

About

3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region. Point clouds are sparse, noisy, and contain limited semantic information compared with 2D images. These inherent limitations make the 3D visual grounding problem more challenging. In this study, we propose 2D Semantics Assisted Training (SAT) that utilizes 2D image semantics in the training stage to ease point-cloud-language joint representation learning and assist 3D visual grounding. The main idea is to learn auxiliary alignments between rich, clean 2D object representations and the corresponding objects or mentioned entities in 3D scenes. SAT takes 2D object semantics, i.e., object label, image feature, and 2D geometric feature, as the extra input in training but does not require such inputs during inference. By effectively utilizing 2D semantics in training, our approach boosts the accuracy on the Nr3D dataset from 37.7% to 49.2%, which significantly surpasses the non-SAT baseline with the identical network architecture and inference input. Our approach outperforms the state of the art by large margins on multiple 3D visual grounding datasets, i.e., +10.4% absolute accuracy on Nr3D, +9.9% on Sr3D, and +5.6% on ScanRef.

Zhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo Luo• 2021

Related benchmarks

TaskDatasetResultRank
3D Visual GroundingScanRefer (val)
Overall Accuracy @ IoU 0.5030.14
155
3D Visual GroundingNr3D (test)
Overall Success Rate56.5
88
3D Visual GroundingNr3D
Overall Success Rate56.5
74
3D Visual GroundingSr3D (test)
Overall Accuracy57.9
73
3D Visual GroundingScanRefer--
23
3D Visual GroundingScanRefer (test)--
21
3D referring expression comprehensionSR3D ReferIt3D (test)
Overall Accuracy57.9
11
3D Object GroundingScanRefer detected proposals v1 (val)
Unique Acc@0.2573.21
10
3D Visual GroundingSr3D
Overall Accuracy57.9
7
3D Dense CaptioningScanRefer Oracle DC
CIDEr80.13
7
Showing 10 of 12 rows

Other info

Follow for update