Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Weakly Supervised Temporal Adjacent Network for Language Grounding

About

Temporal language grounding (TLG) is a fundamental and challenging problem for vision and language understanding. Existing methods mainly focus on fully supervised setting with temporal boundary labels for training, which, however, suffers expensive cost of annotation. In this work, we are dedicated to weakly supervised TLG, where multiple description sentences are given to an untrimmed video without temporal boundary labels. In this task, it is critical to learn a strong cross-modal semantic alignment between sentence semantics and visual content. To this end, we introduce a novel weakly supervised temporal adjacent network (WSTAN) for temporal language grounding. Specifically, WSTAN learns cross-modal semantic alignment by exploiting temporal adjacent network in a multiple instance learning (MIL) paradigm, with a whole description paragraph as input. Moreover, we integrate a complementary branch into the framework, which explicitly refines the predictions with pseudo supervision from the MIL stage. An additional self-discriminating loss is devised on both the MIL branch and the complementary branch, aiming to enhance semantic discrimination by self-supervising. Extensive experiments are conducted on three widely used benchmark datasets, \emph{i.e.}, ActivityNet-Captions, Charades-STA, and DiDeMo, and the results demonstrate the effectiveness of our approach.

Yuechen Wang, Jiajun Deng, Wengang Zhou, Houqiang Li• 2021

Related benchmarks

TaskDatasetResultRank
Temporal GroundingActivityNet Captions
Recall@1 (IoU=0.5)30.01
85
Video Moment RetrievalCharades-STA
R1@0.529.35
57
Video Temporal GroundingCharades-STA
R@1 (IoU=0.5)29.35
48
Video Moment RetrievalActivityNet Captions
R@1 (IoU=0.5)12.38
20
Temporal Sentence GroundingActivityNet Captions v1.3 (test)
Recall (IoU=0.3)52.45
16
Temporal Sentence GroundingCharades-STA (test)
IoU@0.529.35
16
Video Temporal GroundingActivityNet Caption
R@1 (IoU=0.5)30.01
16
Video Moment RetrievalCharades
Rank@1 (IoU=0.3)43.39
16
Video Moment RetrievalActivityNet
Rank@1 (IoU=0.3)52.45
15
Video Moment RetrievalActivityNet Captions
Recall@1 (IoU=0.5)30.01
13
Showing 10 of 15 rows

Other info

Follow for update