Dense Contrastive Learning for Self-Supervised Visual Pre-Training

About

To date, most existing self-supervised learning methods are designed and optimized for image classification. These pre-trained models can be sub-optimal for dense prediction tasks due to the discrepancy between image-level prediction and pixel-level prediction. To fill this gap, we aim to design an effective, dense self-supervised learning method that directly works at the level of pixels (or local features) by taking into account the correspondence between local features. We present dense contrastive learning, which implements self-supervised learning by optimizing a pairwise contrastive (dis)similarity loss at the pixel level between two views of input images. Compared to the baseline method MoCo-v2, our method introduces negligible computation overhead (only <1% slower), but demonstrates consistently superior performance when transferring to downstream dense prediction tasks including object detection, semantic segmentation and instance segmentation; and outperforms the state-of-the-art methods by a large margin. Specifically, over the strong MoCo-v2 baseline, our method achieves significant improvements of 2.0% AP on PASCAL VOC object detection, 1.1% AP on COCO object detection, 0.9% AP on COCO instance segmentation, 3.0% mIoU on PASCAL VOC semantic segmentation and 1.8% mIoU on Cityscapes semantic segmentation. Code is available at: https://git.io/AdelaiDet

Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, Lei Li• 2020

Related benchmarks

Task	Dataset	Result
Semantic segmentation	ADE20K (val)	mIoU37.2	3069
Object Detection	COCO 2017 (val)	AP40.3	2843
Semantic segmentation	PASCAL VOC 2012 (val)	Mean IoU71.6	2204
Image Classification	ImageNet-1k (val)	Top-1 Accuracy63.6	1498
Semantic segmentation	PASCAL VOC 2012 (test)	mIoU69.4	1477
Instance Segmentation	COCO 2017 (val)	APm0.357	1275
Video Object Segmentation	DAVIS 2017 (val)	J mean60.6	1226
Semantic segmentation	ADE20K	mIoU38.1	1028
Image Classification	ImageNet-1k (val)	Top-1 Accuracy63.6	920
Semantic segmentation	Cityscapes	mIoU76.2	668

Showing 10 of 63 rows

Other info

Follow for update

@wizwand_team Discord