Noisy Test-Time Adaptation in Vision-Language Models

About

Test-time adaptation (TTA) aims to address distribution shifts between source and target data by relying solely on target data during testing. In open-world scenarios, models often encounter noisy samples, i.e., samples outside the in-distribution (ID) label space. Leveraging the zero-shot capability of pre-trained vision-language models (VLMs), this paper introduces Zero-Shot Noisy TTA (ZS-NTTA), focusing on adapting the model to target data with noisy samples during test-time in a zero-shot manner. We find existing TTA methods underperform under ZS-NTTA, often lagging behind even the frozen model. We conduct comprehensive experiments to analyze this phenomenon, revealing that the negative impact of unfiltered noisy data outweighs the benefits of clean data during model updating. Also, adapting a classifier for ID classification and noise detection hampers both sub-tasks. Built on this, we propose a framework that decouples the classifier and detector, focusing on developing an individual detector while keeping the classifier frozen. Technically, we introduce the Adaptive Noise Detector (AdaND), which utilizes the frozen model's outputs as pseudo-labels to train a noise detector. To handle clean data streams, we further inject Gaussian noise during adaptation, preventing the detector from misclassifying clean samples as noisy. Beyond the ZS-NTTA, AdaND can also improve the zero-shot out-of-distribution (ZS-OOD) detection ability of VLMs. Experiments show that AdaND outperforms in both ZS-NTTA and ZS-OOD detection. On ImageNet, AdaND achieves a notable improvement of $8.32\%$ in harmonic mean accuracy ($\text{Acc}_\text{H}$) for ZS-NTTA and $9.40\%$ in FPR95 for ZS-OOD detection, compared to SOTA methods. Importantly, AdaND is computationally efficient and comparable to the model-frozen method. The code is publicly available at: https://github.com/tmlr-group/ZS-NTTA.

Chentao Cao, Zhun Zhong, Zhanke Zhou, Tongliang Liu, Yang Liu, Kun Zhang, Bo Han• 2025

Related benchmarks

Task	Dataset	Result
Out-of-Distribution Detection	SUN OOD with ImageNet-1k In-distribution (test)	FPR@9517.08	247
Out-of-Distribution Detection	ImageNet-1k ID iNaturalist OOD	FPR954.19	132
OOD Detection	ImageNet-1k ID Average OOD	AUROC0.9558	92
OOD Detection	iNaturalist (OOD) / ImageNet-1k (ID) 1.0 (test)	FPR951.91	90
Out-of-Distribution Detection	CIFAR10 (ID) vs SVHN (OOD)	AUROC99.97	81
OOD Detection	ImageNet SUN	FPR@9526.95	70
Out-of-Distribution Detection	ImageNet-1k (ID) vs Textures (OOD) 1.0 (test)	AUC93.01	64
Out-of-Distribution Detection	ImageNet Far-OOD	AUROC96	58
Out-of-Distribution Detection	CIFAR-10 In-Dist Texture Out-Dist	AUROC99.63	57
OOD Detection	CIFAR-10 (In-distribution) vs LSUN-R (Out-of-distribution)	FPR950.6	50

Showing 10 of 46 rows

Other info

Follow for update

@wizwand_team Discord