Mean Shift Mask Transformer for Unseen Object Instance Segmentation

About

Segmenting unseen objects from images is a critical perception skill that a robot needs to acquire. In robot manipulation, it can facilitate a robot to grasp and manipulate unseen objects. Mean shift clustering is a widely used method for image segmentation tasks. However, the traditional mean shift clustering algorithm is not differentiable, making it difficult to integrate it into an end-to-end neural network training framework. In this work, we propose the Mean Shift Mask Transformer (MSMFormer), a new transformer architecture that simulates the von Mises-Fisher (vMF) mean shift clustering algorithm, allowing for the joint training and inference of both the feature extractor and the clustering. Its central component is a hypersphere attention mechanism, which updates object queries on a hypersphere. To illustrate the effectiveness of our method, we apply MSMFormer to unseen object instance segmentation. Our experiments show that MSMFormer achieves competitive performance compared to state-of-the-art methods for unseen object instance segmentation. The project page, appendix, video, and code are available at https://irvlutd.github.io/MSMFormer

Yangxiao Lu, Yuqiao Chen, Nicholas Ruozzi, Yu Xiang• 2022

Related benchmarks

Task	Dataset	Result
Instance Segmentation	OCID RGB only (test)	AP5073.9	9
Robot Grasping	SceneReplica	Grasping Success Rate71	4
Robot Pick-and-Place	SceneReplica	Pick-and-Place Success65	4

Showing 3 of 3 rows

Other info

Code

Follow for update

@wizwand_team Discord