On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?

About

The development of large vision-language models, notably CLIP, has catalyzed research into effective adaptation techniques, with a particular focus on soft prompt tuning. Conjointly, test-time augmentation, which utilizes multiple augmented views of a single image to enhance zero-shot generalization, is emerging as a significant area of interest. This has predominantly directed research efforts toward test-time prompt tuning. In contrast, we introduce a robust MeanShift for Test-time Augmentation (MTA), which surpasses prompt-based methods without requiring this intensive training procedure. This positions MTA as an ideal solution for both standalone and API-based applications. Additionally, our method does not rely on ad hoc rules (e.g., confidence threshold) used in some previous test-time augmentation techniques to filter the augmented views. Instead, MTA incorporates a quality assessment variable for each view directly into its optimization process, termed as the inlierness score. This score is jointly optimized with a density mode seeking process, leading to an efficient training- and hyperparameter-free approach. We extensively benchmark our method on 15 datasets and demonstrate MTA's superiority and computational efficiency. Deployed easily as plug-and-play module on top of zero-shot models and state-of-the-art few-shot methods, MTA shows systematic and consistent improvements.

Maxime Zanella, Ismail Ben Ayed• 2024

Related benchmarks

Task	Dataset	Result
Image Classification	DTD	Accuracy45.9	599
Image Classification	ImageNet-R	Top-1 Acc77	581
Image Classification	Flowers102	Accuracy68.1	558
Image Classification	Food101	Accuracy85	457
Image Classification	SUN397	Accuracy66.7	450
Image Classification	StanfordCars	Accuracy68.5	384
Image Classification	Aircraft	Accuracy25.2	340
Fine-grained visual classification	FGVC-Aircraft (test)	Top-1 Acc25.32	312
Image Classification	Pets	Accuracy88.2	308
Fine-grained Image Classification	Stanford Cars	Accuracy67.7	284

Showing 10 of 100 rows

...

Other info

Code

Follow for update

@wizwand_team Discord