Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

About

We propose a label-free approach to adapt powerful but generic vision foundation models to specialized scientific domains. Standard supervised fine-tuning is often ill-suited to these settings: labels are scarce, and task-specific training can collapse the model's generality and hurt robustness. We instead leverage metadata to adapt representations to new domains in a self-supervised manner. Our method, FINO, combines a standard self-supervised objective with flexible metadata guidance that handles both highly granular discrete metadata and continuous metadata. It encourages the representation to preserve informative factors while suppressing spurious ones. Across subcellular fluorescence microscopy, Earth observation, wildlife monitoring, and medical imaging, FINO consistently outperforms standard unsupervised domain adaptation and fully supervised adaptation. It also exceeds highly-specialized domain-specific state of the art, while using no task labels for backbone adaptation and only lightweight probes for supervision.

Elouan Gard\`es, Seung Eun Yi, Kartik Ahuja, Th\'eo Moutakanni, Huy V. Vo, Piotr Bojanowski, Wolfgang M. Pernice, Lo\"ic Landrieu, Camille Couprie• 2026

Related benchmarks

TaskDatasetResultRank
Medical Image ClassificationCheXpert
AUC88
39
Image ClassificationiWildCam OOD (test)
F1-macro53.1
19
Protein LocalizationHPA (test)
F1 Score61.2
9
Satellite Image ClassificationfMoW
WGA52.9
9
Clinical predictionMIMIC
AUROC81.8
7
Subcellular Localization ClusteringOpenCell
ARI0.588
6
Subcellular Localization ClassificationOpenCell (test)
F1 Score78.3
4
Semantic segmentation transferFLAIRHub
mIoU62.3
3
Showing 8 of 8 rows

Other info

Follow for update