Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

About

Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal reasoning pipelines. However, recent advances in generative modeling have shown that such features can often be inverted, enabling realistic reconstructions of the underlying image and raising significant privacy risks. We revisit this problem through the lens of reconstruction and propose TrustCLIP, a reconstruction-driven framework that treats a feature-conditioned generator as an explicit privacy adversary. TrustCLIP learns a projection between encoder features and downstream modules that is explicitly optimized to degrade the reconstructions produced by generative attackers while retaining the necessary signals for downstream tasks. Unlike prior defenses that rely on discriminative privacy metrics, TrustCLIP directly optimizes against a generative reconstruction attacker, targeting a threat not captured by standard evaluation protocols. We demonstrate its effectiveness in both conventional classification and multimodal large language model pipelines. Across these settings, TrustCLIP consistently reduces the fidelity of generative inversions while maintaining downstream task performance. Project page: https://atnikos.github.io/trustclip/

Nikos Athanasiou, Ilya A. Petrov, Angela Yao, Shugao Ma, Eric Sauser, Edoardo Remelli, Shreyas Hampali, Johannes Sch\"onberger, Fadime Sener, Bugra Tekin• 2026

Related benchmarks

TaskDatasetResultRank
Multimodal EvaluationMM-Vet--
249
Object Hallucination EvaluationPOPE
Accuracy (POPE)86.1
137
Multi-modal Perception EvaluationMME Perception
Perception Score1.39e+3
43
Multimodal EvaluationLLaVA-Bench-Wild (LLaVA-W)
Overall Score65.3
43
Visual Question AnsweringScienceQA (SQAI)
Accuracy62.1
42
Visual Question AnsweringTextVQA
VQAText Score55
21
Visual Question AnsweringVQA v2
VQAv2 Score76.3
16
Multimodal EvaluationSEED-Bench Image
SEEDI Score63
7
Image ReconstructionReconstruction Quality under IP-Adapter Attack
PSNR10.42
3
Showing 9 of 9 rows

Other info

Follow for update