Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

About

Brain MRI poses a fundamental challenge for machine learning: models must learn from high-dimensional 3D data spanning multiple co-registered modalities, despite the limited sample sizes typical of neuroimaging studies relative to the diversity in anatomy, pathology, and acquisition conditions. While multimodal imaging provides complementary information critical for clinical interpretation, effectively integrating these signals remains difficult. We propose Multimodal Intra- and Cross-Context Vision Transformer (MICViT), a 3D vision transformer that explicitly models both modality-specific representations and cross-modal interactions across local and global contexts. Concretely, MICViT combines four attention mechanisms: modality-specific local and global attention for intra-modal feature learning, and cross-modal local and global attention to capture interactions between modalities. We evaluate MICViT on brain age prediction across three heterogeneous datasets (UK Biobank, n=41,404; SOOP, n=1,062; Cam-CAN, n=613) using multiple MRI modalities (e.g. T1, FLAIR, DWI, SWI). MICViT consistently outperforms state-of-the-art CNN and transformer baselines in 3D settings. Notably, it benefits more strongly from multimodal inputs, yielding larger performance gains as additional modalities are incorporated. These results demonstrate that explicitly modeling intra- and cross-modal interactions is key to unlocking the full potential of multimodal brain MRI, highlighting a promising direction for representation learning in neuroimaging.

Minh Duc Do, Tillmann Rheude, Noel Kronenberg, Roland Eils, Benjamin Wild• 2026

Related benchmarks

TaskDatasetResultRank
Age regressionUKB
MAE2.816
52
Brain Age Predictioncam-CAN
MAE (years)6.43
11
Brain Age PredictionSOOP
MAE6.67
8
Brain Age PredictionUKBB n=41,404
MAE2.82
8
Showing 4 of 4 rows

Other info

Follow for update