Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning

About

Monocular video depth estimation requires temporal consistency, geometric accuracy, and generalization across diverse scenarios, yet existing methods struggle to achieve all three simultaneously. Discriminative models excel at per-frame accuracy but suffer from temporal drift due to limited context windows, while generative methods improve consistency and generalization at the cost of extensive training data (10M+ samples) and lack of geometric precision. In response to these issues, we introduce \textbf{ICDepth}, a framework that adapts pre-trained text-to-video diffusion transformers for video depth estimation via In-Context Conditioning (ICC), leveraging their rich spatial-temporal priors. To address key challenges in transferring ICC from generation to dense prediction, we propose: (1)~\textbf{SAND-Attention}, which ensures precise spatial-temporal alignment via shared RoPE and enforces unidirectional attention to prevent noise contamination; (2)~\textbf{SRFM}, which injects DINOv2 semantic and resolution priors to enhance geometric precision. ICDepth achieves state-of-the-art results on multiple benchmarks with remarkable data efficiency, trained on only 0.8M frames ($6$--$13\times$ less than competing generative methods), while demonstrating strong zero-shot generalization to diverse domains.

Xuanhua He, Jiaxin Xie, Mingzhe Zheng, Qifeng Chen• 2026

Related benchmarks

TaskDatasetResultRank
Monocular Depth EstimationSintel
Abs Rel0.186
142
Depth EstimationKITTI 110 frames
AbsRel6.1
75
Video Depth EstimationBonn 110 frames
AbsRel5.3
69
Video Depth EstimationScannet 90 frames
AbsRel0.076
28
Video Depth EstimationSintel ~50 frames
AbsRel25
15
Video Depth EstimationScanNet++ (500-frame sequences)
TAE2.61
3
Showing 6 of 6 rows

Other info

Follow for update