Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CILP-FGDI: Exploiting Vision-Language Model for Generalizable Person Re-Identification

About

The Visual Language Model, known for its robust cross-modal capabilities, has been extensively applied in various computer vision tasks. In this paper, we explore the use of CLIP (Contrastive Language-Image Pretraining), a vision-language model pretrained on large-scale image-text pairs to align visual and textual features, for acquiring fine-grained and domain-invariant representations in generalizable person re-identification. The adaptation of CLIP to the task presents two primary challenges: learning more fine-grained features to enhance discriminative ability, and learning more domain-invariant features to improve the model's generalization capabilities. To mitigate the first challenge thereby enhance the ability to learn fine-grained features, a three-stage strategy is proposed to boost the accuracy of text descriptions. Initially, the image encoder is trained to effectively adapt to person re-identification tasks. In the second stage, the features extracted by the image encoder are used to generate textual descriptions (i.e., prompts) for each image. Finally, the text encoder with the learned prompts is employed to guide the training of the final image encoder. To enhance the model's generalization capabilities to unseen domains, a bidirectional guiding method is introduced to learn domain-invariant image features. Specifically, domain-invariant and domain-relevant prompts are generated, and both positive (pulling together image features and domain-invariant prompts) and negative (pushing apart image features and domain-relevant prompts) views are used to train the image encoder. Collectively, these strategies contribute to the development of an innovative CLIP-based framework for learning fine-grained generalized features in person re-identification.

Huazhong Zhao, Lei Qi, Xin Geng• 2025

Related benchmarks

TaskDatasetResultRank
Person Re-IdentificationMSMT17 (test)
Rank-1 Acc61.4
499
Person Re-IdentificationMarket-1501 (test)
Rank-191.3
397
Person Re-IdentificationCUHK03
R135.6
284
Person Re-IdentificationMarket-1501 to DukeMTMC-reID (test)
Rank-164.1
191
Person Re-IdentificationDukeMTMC-reID to Market-1501 (test)
Rank-1 Acc75.7
138
Person Re-IdentificationMSMT17 source: DukeMTMC-reID (test)
Rank-1 Acc66.9
97
Person Re-IdentificationCUHK03 NP (test)
Rank-150.1
69
Person Re-IdentificationAverage (CUHK03-NP, Market-1501, MSMT17)
Rank-167.6
55
Person Re-IdentificationMSMT17 to Market-1501
mAP44.7
46
Person Re-IdentificationMSMT17 MS
mAP28.8
39
Showing 10 of 14 rows

Other info

Follow for update