Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models

About

Vision-Language Models (VLMs) rely heavily on pretrained vision encoders to support downstream tasks such as image captioning, visual question answering, and zero-shot classification. Despite their strong performance, these encoders remain highly vulnerable to imperceptible adversarial perturbations, which can severely degrade both robustness and semantic quality in multimodal reasoning. In this work, we introduce Sim-CLIP, an unsupervised adversarial fine-tuning framework that enhances the robustness of the CLIP vision encoder while preserving overall semantic representations. Sim-CLIP adopts a Siamese training architecture with a cosine similarity objective and a symmetric stop-gradient mechanism to enforce semantic alignment between clean and adversarial views. This design avoids large-batch contrastive learning and additional momentum encoders, enabling robust training with low computational overhead. We evaluate Sim-CLIP across multiple Vision-Language Models and tasks under both targeted and untargeted adversarial attacks. Experimental results demonstrate that Sim-CLIP consistently outperforms state-of-the-art robust CLIP variants, achieving stronger adversarial robustness while maintaining or improving semantic fidelity. These findings highlight the limitations of existing adversarial defenses and establish Sim-CLIP as an effective and scalable solution for robust vision-language representation learning.

Md Zarif Hossain, Ahmed Imteaj• 2024

Related benchmarks

TaskDatasetResultRank
Hallucination EvaluationPOPE--
281
Visual Question AnsweringVQA v2
Accuracy65.08
257
Visual Question AnsweringVizWiz
Accuracy42.49
193
Image CaptioningCOCO
CIDEr106.8
64
Image CaptioningCOCO
CIDEr Drop (%)7.05
60
Visual Question AnsweringOKVQA
VQA Accuracy53.04
34
Jailbreak AttackHADES
Success Rate (Animal)31.89
23
Image CaptioningCOCO
CIDEr (Clean)125.6
14
Image CaptioningFlickr30K
CIDEr (Clean)80.5
14
Visual Question AnsweringVizWiz
VQA Accuracy (Clean)41.5
14
Showing 10 of 19 rows

Other info

Follow for update