Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

About

Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforcement Learning to generate Chains-of-Thought (CoT), they typically rely solely on outcome-based supervision. Consequently, this often leads to hallucinations, where the reasoning process becomes disconnected from the visual content and the final prediction. Existing attempts to mitigate this by relying on external supervision from larger models or separate reward models are computationally expensive and prone to rigid patterns. To address these challenges, we propose TAR (Temporal Anchor-Constrained Reasoning), a framework that introduces the temporal anchor (T-anchor) as a transparent and auditable checkpoint mechanism. T-anchor enforces progressive refinement within the CoT, compelling the model to continuously ground its intermediate thoughts in visual evidence and iteratively calibrate temporal predictions, thereby significantly enhancing the faithfulness and autonomy of the reasoning process and final accuracy. Furthermore, we introduce a bootstrapping paradigm that automatically harvests high-quality CoT data using only a standard 7B model, eliminating the dependency on ultra-large models. Extensive experiments demonstrate that TAR achieves state-of-the-art performance and generates faithful, autonomous, and progressively refined reasoning traces.

Chaohong Guo, Xun Mo, Yongwei Nie, Fei Ma, Xuemiao Xu, Chengjiang Long• 2025

Related benchmarks

TaskDatasetResultRank
Video Question AnsweringVideoMME
Accuracy54.9
254
Temporal GroundingCharades-STA
mIoU61.1
120
Temporal GroundingActivityNet Captions
Recall@1 (IoU=0.5)39.8
85
Video Question AnsweringMVBench
Accuracy66.1
72
Video Temporal GroundingActivityNet Captions
Recall @ IoU=0.361.5
47
Video Temporal GroundingQVHighlights
R1@0.576.1
44
Temporal Video GroundingCharades-STA
Recall@0.571.4
25
Video highlight detectionQVHighlights
R1@0.576.1
15
Showing 8 of 8 rows

Other info

Follow for update