Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training

About

Masked Autoencoders (MAE) have shown promising performance in self-supervised learning for both 2D and 3D computer vision. However, existing MAE-style methods can only learn from the data of a single modality, i.e., either images or point clouds, which neglect the implicit semantic and geometric correlation between 2D and 3D. In this paper, we explore how the 2D modality can benefit 3D masked autoencoding, and propose Joint-MAE, a 2D-3D joint MAE framework for self-supervised 3D point cloud pre-training. Joint-MAE randomly masks an input 3D point cloud and its projected 2D images, and then reconstructs the masked information of the two modalities. For better cross-modal interaction, we construct our JointMAE by two hierarchical 2D-3D embedding modules, a joint encoder, and a joint decoder with modal-shared and model-specific decoders. On top of this, we further introduce two cross-modal strategies to boost the 3D representation learning, which are local-aligned attention mechanisms for 2D-3D semantic cues, and a cross-reconstruction loss for 2D-3D geometric constraints. By our pre-training paradigm, Joint-MAE achieves superior performance on multiple downstream tasks, e.g., 92.4% accuracy for linear SVM on ModelNet40 and 86.07% accuracy on the hardest split of ScanObjectNN.

Ziyu Guo, Renrui Zhang, Longtian Qiu, Xianzhi Li, Pheng-Ann Heng• 2023

Related benchmarks

Task	Dataset	Result
3D Point Cloud Classification	ModelNet40 (test)	OA94	297
Object Classification	ScanObjectNN OBJ_BG	Accuracy90.94	248
Point Cloud Classification	ModelNet40 (test)	Accuracy94.1	229
Object Classification	ScanObjectNN PB_T50_RS	Accuracy86.07	220
Object Classification	ScanObjectNN OBJ_ONLY	Overall Accuracy88.86	186
Few-shot classification	ModelNet40 10-way 10-shot	Accuracy92.6	117
Few-shot classification	ModelNet40 10-way 20-shot	Accuracy95.1	117
Few-shot classification	ModelNet40 5-way 10-shot	Accuracy96.7	102
Few-shot classification	ModelNet40 5-way 20-shot	Accuracy97.9	102
Point Cloud Classification	ScanObjectNN PB_T50_RS	Overall Accuracy86.07	100

Showing 10 of 35 rows

Other info

Follow for update

@wizwand_team Discord