VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining
About
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Monocular Depth Estimation | NYU v2 (test) | Abs Rel0.082 | 327 | |
| Monocular Depth Estimation | KITTI | AbsRel6 | 33 | |
| 3D Semantic Segmentation | ScanNet200 | mIoU12.1 | 28 | |
| Camera pose estimation | Sintel | ATE0.088 | 16 | |
| Camera pose estimation | TUM-RGBD dynamics | ATE0.09 | 11 | |
| 3D Semantic Segmentation | KITTI Odometry Open-Vocabulary 3D Segmentation Protocol | mIoU24 | 5 | |
| Intrinsic parameter estimation | Sintel | AFE (px)255.5 | 5 | |
| 3D Semantic Segmentation | KITTI | mIoU24 | 2 |