Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging
About
Brain MRIs are routinely acquired as multiple complementary sequences with unique contrast weighting, including T1-weighed imaging (T1w) anatomic and fluid-sensitive T2-weighted (T2w) contrasts. However, methods for learning unified representations across the multitude of MRI contrast mechanisms at health-system scale are lacking. In this study, we introduce Neuro-JEPA, a sparse multimodal neuroimaging foundation model that combines a latent predictive objective with a Mixture-of-Experts architecture to encode brain MRI across core T1w, T2w, and fluid-suppressed FLAIR imaging (FLAIR). We further provide a systematic methodological study of architectural, masking, objective, and sparsity design choices beneficial for robust neuroimaging multimodal representation learning. Neuro-JEPA was pretrained on 1,551,862 scans from 428,647 studies after modality-specific preprocessing with data curation across three core structural brain MRI sequences. We evaluated the learned representations across clinical and research settings, including 25 tasks from three health systems: NYU Langone, NYU Long Island, and Massachusetts General Hospital, and 22 tasks from 12 public datasets, covering unimodal, multimodal and cross-domain evaluation configurations. Across these benchmarks, existing neuroimaging foundation models showed inconsistent gains over a simple convolutional neural network (CNN) baseline, whereas Neuro-JEPA achieved stronger and more consistent performance across all evaluated settings. These results establish a scalable methodological framework for multimodal neuroimaging representation learning and highlight the need for foundation model evaluation protocols that include simple baselines, clinically heterogeneous cohorts and controlled multimodal comparisons.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Brain Age Prediction | OpenBHB (test) | R20.894 | 4 | |
| Diagnosis and prognosis | Public 41 combs | AUROC78.5 | 4 | |
| Diagnosis and prognosis | NYU Langone 30 combs | AUROC0.825 | 4 | |
| Diagnosis and prognosis | NYU Long Island 30 combs | AUROC80.6 | 4 | |
| Diagnosis and prognosis | BIND-MGH 45 combs | AUROC74.1 | 4 | |
| IDH mutation prediction | UCSF-PDGM | Macro AUROC83.2 | 4 | |
| Length-of-Stay Prediction | ICSPR-Stroke | Macro AUROC75.4 | 4 | |
| Multimodal Learning | Public datasets 12 tasks | AUROC0.805 | 4 | |
| Multimodal Learning | BIND-MGH 30 tasks | AUROC76.3 | 4 | |
| Time-to-event prediction | Time-to-Event 6 combs | C-index0.695 | 4 |