Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Knowledge-Preserved Model Tuning in Null-Space for Robust Spatio-Temporal Video Grounding

About

Spatio-Temporal Video Grounding aims to localize object tubes based on textual queries. While recent methods have achieved remarkable success, they mainly focus on high-quality(HQ) inputs, neglecting the widespread presence of low-quality(LQ) videos in real-world scenarios. Although tuning methods like LoRA can adapt to degraded inputs, they inevitably disrupt pre-trained knowledge. To address this, we propose Null-Space Tuning (NST). This framework exploits the geometric property that adding vectors within the null-space of frozen weights to the layer input does not affect the output. Leveraging this, NST injects learnable residuals into input features that can be selectively invisible to the pre-trained backbone. Specifically, NST combines the Quality-Adaptive Unit and Dual-Space Reparameterization to synthesize these residuals by confining components for HQ inputs to the null-space, while directing restoration components for LQ inputs to the non-null space. As the frozen weights eliminate null-space components, we effectively rectify degraded inputs while preserving pre-trained knowledge for HQ inputs. Extensive experiments show that NST outperforms state-of-the-art methods on our Mixed-Quality benchmark.

Haoxuan Chen, Xianqin Liu, Jian-Fang Hu• 2026

Related benchmarks

TaskDatasetResultRank
Spatio-Temporal Video GroundingVidSTG Interrogative Sentences (test)
m_vIoU26.49
47
Video Spatio-Temporal GroundingVidSTG Declarative Sentences Mixed-Quality (test)
Mean Temporal IoU (mIoU)50.81
7
Video Spatio-Temporal GroundingVidSTG HQ and LQ Declarative Sentences v1 (test)
Delta Mean Temporal IoU1.17
6
Video Spatio-Temporal GroundingVidSTG HQ and LQ Interrogative Sentences v1 (test)
Delta Mean Temporal IoU0.71
6
Showing 4 of 4 rows

Other info

Follow for update