Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding

About

Spatio-Temporal video grounding (STVG) focuses on retrieving the spatio-temporal tube of a specific object depicted by a free-form textual expression. Existing approaches mainly treat this complicated task as a parallel frame-grounding problem and thus suffer from two types of inconsistency drawbacks: feature alignment inconsistency and prediction inconsistency. In this paper, we present an end-to-end one-stage framework, termed Spatio-Temporal Consistency-Aware Transformer (STCAT), to alleviate these issues. Specially, we introduce a novel multi-modal template as the global objective to address this task, which explicitly constricts the grounding region and associates the predictions among all video frames. Moreover, to generate the above template under sufficient video-textual perception, an encoder-decoder architecture is proposed for effective global context modeling. Thanks to these critical designs, STCAT enjoys more consistent cross-modal feature alignment and tube prediction without reliance on any pre-trained object detectors. Extensive experiments show that our method outperforms previous state-of-the-arts with clear margins on two challenging video benchmarks (VidSTG and HC-STVG), illustrating the superiority of the proposed framework to better understanding the association between vision and natural language. Code is publicly available at https://github.com/jy0205/STCAT.

Yang Jin, Yongzhi Li, Zehuan Yuan, Yadong Mu• 2022

Related benchmarks

Task	Dataset	Result
Spatio-Temporal Video Grounding	HCSTVG v1 (test)	m_vIoU35.1	42
Spatio-Temporal Video Grounding	VidSTG Interrogative Sentences (test)	m_vIoU28.22	40
Spatio-Temporal Video Grounding	VidSTG Declarative Sentences (test)	m_vIoU33.14	24
Spatio-Temporal Video Grounding	VidSTG Declarative Sentences	m_vIoU33.1	20
Spatio-Temporal Video Grounding	HC-STVG (val)	Mean vIoU31.2	19
Spatio-Temporal Video Grounding	VidSTG Declarative (test)	m_vIoU33.1	14
Spatio-Temporal Video Grounding	HC-STVG v1 (test)	m_vIoU35	14
Action Grounding	Daly (test)	Accuracy55.9	13
Spatio-Temporal Video Grounding	HC-STVG v1	m_vIoU35.1	11
Spatio-Temporal Video Grounding	VidSTG (test)	Declarative m_vIoU33.1	11

Showing 10 of 14 rows

Other info

Code

Follow for update

@wizwand_team Discord