Single-Stage 6D Object Pose Estimation

About

Most recent 6D pose estimation frameworks first rely on a deep network to establish correspondences between 3D object keypoints and 2D image locations and then use a variant of a RANSAC-based Perspective-n-Point (PnP) algorithm. This two-stage process, however, is suboptimal: First, it is not end-to-end trainable. Second, training the deep network relies on a surrogate loss that does not directly reflect the final 6D pose estimation task. In this work, we introduce a deep architecture that directly regresses 6D poses from correspondences. It takes as input a group of candidate correspondences for each 3D keypoint and accounts for the fact that the order of the correspondences within each group is irrelevant, while the order of the groups, that is, of the 3D keypoints, is fixed. Our architecture is generic and can thus be exploited in conjunction with existing correspondence-extraction networks so as to yield single-stage 6D pose estimation frameworks. Our experiments demonstrate that these single-stage frameworks consistently outperform their two-stage counterparts in terms of both accuracy and speed.

Yinlin Hu, Pascal Fua, Wei Wang, Mathieu Salzmann• 2019

Related benchmarks

Task	Dataset	Result
6D Pose Estimation	YCB-Video	--	151
6DoF Pose Estimation	YCB-Video (test)	--	72
6D Object Pose Estimation	OccludedLINEMOD (test)	ADD(S)43.3	45
6D Pose Estimation	YCB-V	--	29
6D Object Pose Estimation	LM-O (test)	--	22
6DoF Pose Estimation	Occlusion Linemod (Part I)	Average Error43.3	16
6DoF Pose Estimation	Linemod (Part II)	Ape Accuracy66.7	4

Showing 7 of 7 rows

Other info

Follow for update

@wizwand_team Discord