Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

About

Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.

Marwane Hariat, David Filliat, Antoine Manzanera• 2026

Related benchmarks

TaskDatasetResultRank
Monocular Depth EstimationNYU v2 (test)
Abs Rel0.082
327
Monocular Depth EstimationKITTI
AbsRel6
33
3D Semantic SegmentationScanNet200
mIoU12.1
28
Camera pose estimationSintel
ATE0.088
16
Camera pose estimationTUM-RGBD dynamics
ATE0.09
11
3D Semantic SegmentationKITTI Odometry Open-Vocabulary 3D Segmentation Protocol
mIoU24
5
Intrinsic parameter estimationSintel
AFE (px)255.5
5
3D Semantic SegmentationKITTI
mIoU24
2
Showing 8 of 8 rows

Other info

Follow for update