Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation Models

About

Foundation models have driven rapid progress in computer vision, yet the two dominant paradigms, vision-language foundation models (VLMs) and vision-only foundation models (VFMs), remain only partially compatible. VLMs offer language-grounded semantic alignment but are often visually coarse, while VFMs learn discriminative perceptual geometry but lack semantic grounding. We propose GPUA (Geometry-Preserving Unsupervised Alignment), a framework that integrates the complementary strengths of VFMs and VLMs. Inspired by cross-lingual alignment, GPUA treats VFM features as a visual language and learns an orthogonal mapping that translates the VFM space into the VLM semantic space, preserving geometry and narrowing the modality gap without labels or model parameter updates. GPUA is task-agnostic and requires only feature-level access to pretrained models. Experiments across diverse benchmarks demonstrate improved cross-model compatibility and strong gains in downstream zero-shot recognition and segmentation with negligible overhead. Code is available at https://github.com/Yuteam14/GPUA

Shuwen Yu, Zhanxuan Hu, Yi Zhao, Yonghang Tai, Huafeng Li• 2026

Related benchmarks

TaskDatasetResultRank
ClassificationCars
Accuracy77.7
571
Semantic segmentationCityscapes
mIoU42
526
Image ClassificationImageNet
Top-1 Accuracy75.4
384
Image ClassificationPets
Accuracy95
320
Image ClassificationImageNet V2 (test)--
232
Semantic segmentationPascal Context 59
mIoU41
217
Image ClassificationImageNet-R (test)
Accuracy77.5
179
Image ClassificationImageNet-A (test)--
177
Semantic segmentationPascal VOC 20
mIoU87.8
151
Image ClassificationCaltech
Accuracy98.1
145
Showing 10 of 21 rows

Other info

Follow for update