Exploring Visual Prompts for Adapting Large-Scale Models

About

We investigate the efficacy of visual prompting to adapt large-scale models in vision. Following the recent approach from prompt tuning and adversarial reprogramming, we learn a single image perturbation such that a frozen model prompted with this perturbation performs a new task. Through comprehensive experiments, we demonstrate that visual prompting is particularly effective for CLIP and robust to distribution shift, achieving performance competitive with standard linear probes. We further analyze properties of the downstream dataset, prompt design, and output transformation in regard to adaptation performance. The surprising effectiveness of visual prompting provides a new perspective on adapting pre-trained models in vision. Code is available at http://hjbahng.github.io/visual_prompting .

Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, Phillip Isola• 2022

Related benchmarks

Task	Dataset	Result
Image Classification	Stanford Cars	--	660
Image Classification	Food-101	Accuracy81.8	570
Image Classification	EuroSAT	Accuracy90.8	569
Image Classification	UCF101	Top-1 Acc74.2	527
Image Classification	RESISC45	Accuracy81.4	472
Image Classification	SVHN (test)	Accuracy60.4	470
Image Classification	Food101	Accuracy78.1	457
Image Classification	SUN397	--	450
Action Recognition	UCF101	Accuracy67.9	433
Image Classification	SUN397	Accuracy67.1	425

Showing 10 of 73 rows

...

Other info

Follow for update

@wizwand_team Discord