Beyond Templates: Revisiting Zero-Shot Remote Sensing through Meta-Prompting
About
Vision-language models (VLMs) have sparked growing interest in zero-shot Earth Observation (EO) downstream tasks, with further gains enabled by remote-sensing-adapted models. We examine this setting across 17 VLM variants and 12 remote sensing (RS) datasets under Meta-Prompting for Visual Recognition (MPVR), and show that zero-shot performance remains highly sensitive to textual design choices, from the meta-prompts used to guide the LLM in generating class descriptions to the descriptions themselves. We explore why semantically rich LLM-generated class descriptions do not translate into consistent gains over simple domain-adapted CLIP-style descriptions. While LLM descriptions are more semantically expressive, they can also introduce noise in the text embedding space, reducing robustness in downstream tasks. We support this observation through a text log-likelihood analysis in the whitened CLIP feature space, comparing LLM-generated and template-based descriptions. Building on this finding, we study query embedding calibration and show that lightweight calibration of the query space consistently yields strong improvements in zero-shot classification and retrieval. Overall, our results provide practical insight into the trade-off between semantic richness and robustness, and identify embedding calibration as a simple and effective tool for improving zero-shot remote sensing performance.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Image Classification | RESISC45 | Accuracy75.4 | 539 | |
| Image Classification | WHU-RS19 | Accuracy95.3 | 104 | |
| Image Classification | AID | Accuracy91.5 | 83 | |
| Classification | EuroSAT | Accuracy69.9 | 75 | |
| Image Classification | PatternNet | Accuracy82.8 | 65 | |
| Remote Sensing Classification | SIRI-WHU | Top-1 Acc72.6 | 62 | |
| Image Classification | RSI-CB128 | Accuracy48.5 | 60 | |
| Image Classification | RSI-CB256 | Accuracy55.7 | 60 | |
| Image Classification | MLRSNet | Accuracy70.7 | 60 | |
| Image Classification | RSC11 | Accuracy78.1 | 60 |