3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy
About
Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We present a systematic comparison of 2D and 3D masked autoencoders (MAE-2D vs. MAE-3D) on volumetric microscopy data. Under matched architectures and training protocols, MAE-3D consistently outperforms 2D max-projection and slice-based variants on downstream single-cell tasks. We further align visual representations with a pretrained protein language model (ESM2) and show that cross-modal supervision yields larger gains for volumetric models. Channel cross-attention and frequency-domain regularization are critical for leveraging 3D spatial context. On protein--protein interaction prediction, our best model achieves a ROC--AUC of 0.86, while on protein localization it reaches an AUC$_{\text{micro}}$ of 0.95 and an F1$_{\text{micro}}$ of 0.74, demonstrating competitive performance on both tasks. Overall, our findings highlight the potential of volumetric modeling and multimodal alignment for representation learning in single-cell microscopy.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Protein Localization | OpenCell (five-fold splits) | mAP62 | 4 |