Do Convnets Learn Correspondence?

About

Convolutional neural nets (convnets) trained from massive labeled datasets have substantially improved the state-of-the-art in image classification and object detection. However, visual understanding requires establishing correspondence on a finer level than object category. Given their large pooling regions and training from whole-image labels, it is not clear that convnets derive their success from an accurate correspondence model which could be used for precise localization. In this paper, we study the effectiveness of convnet activation features for tasks requiring correspondence. We present evidence that convnet features localize at a much finer scale than their receptive field sizes, that they can be used to perform intraclass alignment as well as conventional hand-engineered features, and that they outperform conventional features in keypoint prediction on objects from PASCAL VOC 2011.

Jonathan Long, Ning Zhang, Trevor Darrell• 2014

Related benchmarks

Task	Dataset	Result
2D Keypoint Localization	PASCAL3D+ (test)	Aero Acc53.7	6
Keypoint Localization	PASCAL VOC 2012	Aero Acc53.7	5
Keypoint Prediction	PASCAL VOC 2011	PCK (Aeroplane)50.9	4
Keypoint Transfer	PASCAL VOC 2011 (test)	Aero28.2	3
Keypoint Classification	Pascal3D+ 44 (test)	Accuracy (Aero)0.44	3
Human keypoint localization	PASCAL VOC Person 2011 (val)	PCK (alpha=0.10)47.1	3

Showing 6 of 6 rows

Other info

Follow for update

@wizwand_team Discord