Video Relation Detection via Tracklet based Visual Transformer

About

Video Visual Relation Detection (VidVRD), has received significant attention of our community over recent years. In this paper, we apply the state-of-the-art video object tracklet detection pipeline MEGA and deepSORT to generate tracklet proposals. Then we perform VidVRD in a tracklet-based manner without any pre-cutting operations. Specifically, we design a tracklet-based visual Transformer. It contains a temporal-aware decoder which performs feature interactions between the tracklets and learnable predicate query embeddings, and finally predicts the relations. Experimental results strongly demonstrate the superiority of our method, which outperforms other methods by a large margin on the Video Relation Understanding (VRU) Grand Challenge in ACM Multimedia 2021. Codes are released at https://github.com/Dawn-LX/VidVRD-tracklets.

Kaifeng Gao, Long Chen, Yifeng Huang, Jun Xiao• 2021

Related benchmarks

Task	Dataset	Result	Rank
Visual Relation Tagging	VidOR (val)	P@552.73		14
Visual Relation Detection	VidOR (val)	R@500.0835		13

Showing 2 of 2 rows

Other info

Follow for update

@wizwand_team Discord