Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language Models

About

Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations. Recent test-time adaptation defenses improve robustness by leveraging many augmented views, but this leads to impractical slowdown and a clear robustness-throughput trade-off. To address this challenge, we present Stability and Suitability-guided Test-time Prompt Tuning (SS-TPT), evaluating the quality of each augmented view via two complementary scores: (1) stability, measuring prediction invariance to weak augmentations, and (2) suitability, measuring feature-space density among views. These stability and suitability (SS) scores guide both adaptation and inference through an SS-guided consistency loss and an SS-weighted prediction, amplifying trustworthy views while suppressing corrupted ones. Extensive experiments demonstrate that SS-TPT significantly outperforms prior state-of-the-art methods, achieving superior robustness-throughput trade-offs across diverse datasets and varying numbers of views, thereby demonstrating both strong practicality and generality. Our code is available at https://github.com/sunoh-kim/SS-TPT.

Sunoh Kim, Daeho Um• 2026

Related benchmarks

TaskDatasetResultRank
ClassificationCars--
571
Fine grained classificationEuroSAT
Accuracy34.1
138
Image ClassificationImageNet A
Accuracy26.9
98
Fine grained classificationDTD
Clean Accuracy45
54
Image ClassificationEuroSAT
Top-1 Clean Accuracy22.6
53
Fine-grained Image ClassificationUCF101
Accuracy66
47
Image ClassificationImageNet-S
Accuracy33.2
46
Fine grained classificationPets (test)
Accuracy83.5
42
Image ClassificationUCF101
Accuracy58.3
42
Image ClassificationCaltech101
Top-1 Robust Accuracy81.2
40
Showing 10 of 37 rows

Other info

Follow for update