Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Multi-Modal Contrastive Learning for Implicit Earth Embeddings via Location Tying

About

Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations. To overcome this challenge, self-supervised pre-training is a possible solution, with contrastive learning dominant for location encoders. Those approaches usually align geographic coordinates with just one additional modality. We propose two multimodal contrastive learning architectures: Multimodal Embedding via Location Tying (MELT) and Sequential Alternating Location Training (SALT). These architectures expand this framework beyond two modalities by utilising unpaired geospatial data. Both methods are technically viable and match the performance of the strongest two-modality baseline (SATCLIP) across four downstream tasks. However, increasing the number of modalities does not consistently improve performance, suggesting that the chosen location encoder is the main limitation - the contrastive objective reaches its peak early, regardless of modality diversity or pre-training volume. MELT provides more stable training than SALT and presents a stronger foundation for future scaling.

Jonathan Hecht, Lukas Arzoumanidis, Ziyue Li, Youness Dehbi• 2026

Related benchmarks

TaskDatasetResultRank
RegressionElevation
0.8
59
ClassificationCountry
Accuracy89.6
46
ClassificationBiome
Accuracy70.1
37
Biome ClassificationBiome 3125 samples
Accuracy82.5
22
Population Density RegressionPopulation 1024 samples
R20.63
22
Biome ClassificationBiome 1024 samples
Accuracy77.4
22
Country ClassificationCountry 243 samples
Accuracy69.3
22
Country ClassificationCountry 1024 samples
Accuracy83.2
22
Elevation RegressionElevation 1024 samples
R273
22
Population Density RegressionPopulation 243 samples
R20.57
22
Showing 10 of 11 rows

Other info

Follow for update