Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

About

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive segmentation losses, which overlooks the geometric consistency across frames and leads to weak spatial understanding. In this paper, we propose Geometry-enhanced Language-guided Video segmentation (GeoLaV), a two-stage framework that distills 3D geometric knowledge from images to enhance text-driven video segmentation. In the first stage, we perform monocular geometry pretraining with monocular novel-view synthesis, enabling the model to acquire geometry-consistent visual representations via spatial alignment on large-scale single-image datasets. In the second stage, we introduce geometry-aware distillation and fine-tune the model on video segmentation datasets, transferring 3D structural knowledge from a general 3D prior model. This process reinforces 3D awareness and improves both spatiotemporal coherence and language grounding in segmentation. Extensive experiments show that our method using only image segmentation data already provides notable zero-shot generalization in RVOS. When combined with geometry-aware distillation for fine-tuning on videos, our method achieves state-of-the-art performance across multiple RVOS benchmarks. The code is available at https://github.com/Tony1882880/GeoLaV.

Tianyu Zhu, Yingping Liang, Hesong Li, Ying Fu• 2026

Related benchmarks

TaskDatasetResultRank
Referring Video Object SegmentationRef-DAVIS 17
J&F Score72.5
165
Referring Video Object SegmentationRef-YouTube-VOS
J&F70.5
143
Referring Video Object SegmentationMeViS
J&F Score50
31
Referring Video Object SegmentationRef-YouTube-VOS 45 (test)
J&F Score47
8
Referring Video Object SegmentationMeViS 6 (test)
J&F Score31.6
8
Showing 5 of 5 rows

Other info

Follow for update