Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Repurposing CLIP to Localize at Pixel Level

About

Large-scale Vision-Language Models like CLIP have demonstrated impressive open-set localization capabilities at the image level. However, adapting this capability to pixel-level dense prediction poses challenges due to global feature biases. In this paper, we introduce CLIPix, a simple yet effective framework that repurposes CLIP to perform pixel-level localization. By tracing back CLIP's classification process, CLIPix identifies object-specific attentive regions and repurposes them as pixel-level localization cues. To address noise introduced by global biases, we propose a Noise-Resistant Correction strategy, refining these cues for more precise segmentation. Additionally, we introduce a Localization Embedding strategy to integrate both localization and enriched detail information, enabling accurate, high-resolution segmentation. Our approach preserves CLIP's generalization strength and unlocks its potential for segmenting arbitrary objects. Extensive experiments on the PASCAL and COCO datasets demonstrate that CLIPix achieves state-of-the-art performance, underscoring its effectiveness.

Jiaxiang Fang, Shiqiang Ma, Jing Wang, Siyu Chen, Fei Guo, Shengfeng He• 2026

Related benchmarks

TaskDatasetResultRank
Semantic segmentationCOCO-20i
mIoU (Mean)61.8
187
Semantic segmentationPASCAL-5i
Mean mIoU80.7
149
Semantic segmentationCOCO-20i (test)--
70
Open Vocabulary Semantic SegmentationPASCAL Context 59 (val)
mIoU37.4
57
Open Vocabulary Semantic SegmentationCityscapes (val)
mIoU37.1
56
Open Vocabulary Semantic SegmentationPASCAL VOC-20 (val)
mIoU87.8
23
Open Vocabulary Semantic SegmentationADE20K (val)
mIoU19.1
19
Showing 7 of 7 rows

Other info

Follow for update