Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

About

Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding. In this paper, we propose UniScene3D, a transformer-based framework that learns unified 3D scene representations from multi-view RGB-Pointmap inputs by leveraging the priors of a pretrained 2D foundation model. For robust RGB-Pointmap representation learning, we introduce cross-view geometric alignment and grounded view alignment to enforce geometric and semantic consistency across views. Extensive low-shot and task-specific fine-tuning on viewpoint grounding, scene retrieval, scene classification, and 3D visual question answering demonstrates state-of-the-art performance. These results establish UniScene3D as an effective framework for unified 3D scene understanding. Project page: https://yebulabula.github.io/UniScene3D/

Ye Mao, Weixun Luo, Ranran Huang, Junpeng Jing, Krystian Mikolajczyk• 2026

Related benchmarks

TaskDatasetResultRank
3D Visual Question AnsweringSQA3D
EM@152.5
30
3D Visual Question AnsweringScanQA
EM@123.2
17
3D Visual Question AnsweringHypo3D
EM@135.2
14
Scene RetrievalScanRefer (n=5)
Recall@122.4
8
Scene RetrievalScanRefer (n=10)
Recall@133.4
8
Scene RetrievalNr3D n=5
R@119.7
8
Scene RetrievalNr3D n=10
R@130.7
8
Scene RetrievalSr3D n=5
R@13
8
Scene RetrievalSr3D (n=10)
R@14.6
8
Scene type classificationScanNet v2 (test)
Accuracy (0-shot)70.7
8
Showing 10 of 14 rows

Other info

Follow for update