Multi-Modal Contrastive Learning for Implicit Earth Embeddings via Location Tying
About
Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations. To overcome this challenge, self-supervised pre-training is a possible solution, with contrastive learning dominant for location encoders. Those approaches usually align geographic coordinates with just one additional modality. We propose two multimodal contrastive learning architectures: Multimodal Embedding via Location Tying (MELT) and Sequential Alternating Location Training (SALT). These architectures expand this framework beyond two modalities by utilising unpaired geospatial data. Both methods are technically viable and match the performance of the strongest two-modality baseline (SATCLIP) across four downstream tasks. However, increasing the number of modalities does not consistently improve performance, suggesting that the chosen location encoder is the main limitation - the contrastive objective reaches its peak early, regardless of modality diversity or pre-training volume. MELT provides more stable training than SALT and presents a stronger foundation for future scaling.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Regression | Elevation | R²0.8 | 59 | |
| Classification | Country | Accuracy89.6 | 46 | |
| Classification | Biome | Accuracy70.1 | 37 | |
| Biome Classification | Biome 3125 samples | Accuracy82.5 | 22 | |
| Population Density Regression | Population 1024 samples | R20.63 | 22 | |
| Biome Classification | Biome 1024 samples | Accuracy77.4 | 22 | |
| Country Classification | Country 243 samples | Accuracy69.3 | 22 | |
| Country Classification | Country 1024 samples | Accuracy83.2 | 22 | |
| Elevation Regression | Elevation 1024 samples | R273 | 22 | |
| Population Density Regression | Population 243 samples | R20.57 | 22 |