CILP-FGDI: Exploiting Vision-Language Model for Generalizable Person Re-Identification

About

The Visual Language Model, known for its robust cross-modal capabilities, has been extensively applied in various computer vision tasks. In this paper, we explore the use of CLIP (Contrastive Language-Image Pretraining), a vision-language model pretrained on large-scale image-text pairs to align visual and textual features, for acquiring fine-grained and domain-invariant representations in generalizable person re-identification. The adaptation of CLIP to the task presents two primary challenges: learning more fine-grained features to enhance discriminative ability, and learning more domain-invariant features to improve the model's generalization capabilities. To mitigate the first challenge thereby enhance the ability to learn fine-grained features, a three-stage strategy is proposed to boost the accuracy of text descriptions. Initially, the image encoder is trained to effectively adapt to person re-identification tasks. In the second stage, the features extracted by the image encoder are used to generate textual descriptions (i.e., prompts) for each image. Finally, the text encoder with the learned prompts is employed to guide the training of the final image encoder. To enhance the model's generalization capabilities to unseen domains, a bidirectional guiding method is introduced to learn domain-invariant image features. Specifically, domain-invariant and domain-relevant prompts are generated, and both positive (pulling together image features and domain-invariant prompts) and negative (pushing apart image features and domain-relevant prompts) views are used to train the image encoder. Collectively, these strategies contribute to the development of an innovative CLIP-based framework for learning fine-grained generalized features in person re-identification.

Huazhong Zhao, Lei Qi, Xin Geng• 2025

Related benchmarks

Task	Dataset	Result
Person Re-Identification	MSMT17 (test)	Rank-1 Acc61.4	517
Person Re-Identification	Market-1501 (test)	Rank-191.3	417
Person Re-Identification	CUHK03	R135.6	322
Person Re-Identification	Market-1501 to DukeMTMC-reID (test)	Rank-164.1	191
Person Re-Identification	DukeMTMC-reID to Market-1501 (test)	Rank-1 Acc75.7	138
Person Re-Identification	MSMT17 source: DukeMTMC-reID (test)	Rank-1 Acc66.9	97
Person Re-Identification	CUHK03 NP (test)	Rank-150.1	69
Person Re-Identification	Average (CUHK03-NP, Market-1501, MSMT17)	Rank-167.6	55
Person Re-Identification	Market-1501 -> MSMT17 (M→MS) (test)	Rank-147	48
Person Re-Identification	MSMT17 to Market-1501	mAP44.7	46

Showing 10 of 19 rows

Other info

Follow for update

@wizwand_team Discord