Is ImageNet worth 1 video? Learning strong image encoders from 1 long unlabelled video

About

Self-supervised learning has unlocked the potential of scaling up pretraining to billions of images, since annotation is unnecessary. But are we making the best use of data? How more economical can we be? In this work, we attempt to answer this question by making two contributions. First, we investigate first-person videos and introduce a "Walking Tours" dataset. These videos are high-resolution, hours-long, captured in a single uninterrupted take, depicting a large number of objects and actions with natural scene transitions. They are unlabeled and uncurated, thus realistic for self-supervision and comparable with human learning. Second, we introduce a novel self-supervised image pretraining method tailored for learning from continuous videos. Existing methods typically adapt image-based pretraining approaches to incorporate more frames. Instead, we advocate a "tracking to learn to recognize" approach. Our method called DoRA, leads to attention maps that Discover and tRAck objects over time in an end-to-end manner, using transformer cross-attention. We derive multiple views from the tracks and use them in a classical self-supervised distillation loss. Using our novel approach, a single Walking Tours video remarkably becomes a strong competitor to ImageNet for several image and video downstream tasks.

Shashanka Venkataramanan, Mamshad Nayeem Rizve, Jo\~ao Carreira, Yuki M. Asano, Yannis Avrithis• 2023

Related benchmarks

Task	Dataset	Result
Object Detection	COCO 2017 (val)	AP23.2	2930
Instance Segmentation	COCO 2017 (val)	--	1304
Video Object Segmentation	DAVIS 2017 (val)	J mean51.9	1251
Image Classification	ImageNet-1K	--	600
Object Tracking	LaSoT	AUC61.7	519
Visual Object Tracking	TrackingNet (test)	Normalized Precision (Pnorm)82.5	502
Visual Object Tracking	GOT-10k (test)	Average Overlap63.8	461
Semantic segmentation	ADE20K	mIoU21.6	90
Video Object Segmentation	DAVIS 2017	Jaccard Index (J)51.9	82
Image Classification	ImageNet-1k (val)	Accuracy34.8	64

Showing 10 of 17 rows

Other info

Follow for update

@wizwand_team Discord