Exploring Set Similarity for Dense Self-supervised Representation Learning

About

By considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue, in this paper, we propose to explore \textbf{set} \textbf{sim}ilarity (SetSim) for dense self-supervised representation learning. We generalize pixel-wise similarity learning to set-wise one to improve the robustness because sets contain more semantic and structure information. Specifically, by resorting to attentional features of views, we establish corresponding sets, thus filtering out noisy backgrounds that may cause incorrect correspondences. Meanwhile, these attentional features can keep the coherence of the same image across different views to alleviate semantic inconsistency. We further search the cross-view nearest neighbours of sets and employ the structured neighbourhood information to enhance the robustness. Empirical evaluations demonstrate that SetSim is superior to state-of-the-art methods on object detection, keypoint detection, instance segmentation, and semantic segmentation.

Zhaoqing Wang, Qiang Li, Guoxin Zhang, Pengfei Wan, Wen Zheng, Nannan Wang, Mingming Gong, Tongliang Liu• 2021

Related benchmarks

Task	Dataset	Result
Object Detection	COCO 2017 (val)	--	2930
Instance Segmentation	COCO 2017 (val)	APm0.364	1304
Semantic segmentation	ADE20K	mIoU38.6	1028
Semantic segmentation	Cityscapes	mIoU77	674
Semantic segmentation	Pascal VOC	mIoU0.709	295
Instance Segmentation	Cityscapes (val)	--	247
Object Detection	PASCAL VOC 2007+2012 (test)	--	100
Keypoint Detection	MS-COCO 2017 (val)	AP66.7	44

Showing 8 of 8 rows

Other info

Follow for update

@wizwand_team Discord