Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

About

Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On protein--protein interaction prediction, our best model achieves a ROC--AUC of 0.86, while on protein localization it reaches an AUC$_{\text{micro}}$ of 0.95 and an F1$_{\text{micro}}$ of 0.74, demonstrating competitive performance on both tasks. Overall, our findings highlight the potential of volumetric modeling and multimodal alignment for representation learning in single-cell microscopy.

Amirhossein Kardoost, Lion Gleiter, Tingying Peng, Carsten Marr• 2026

Related benchmarks

TaskDatasetResultRank
Protein LocalizationOpenCell (five-fold splits)
mAP62
4
Showing 1 of 1 rows

Other info

Follow for update