Mixed-Modality Dual Face-Hair Retrieval
About
We introduce Dual Face-Hair Retrieval (DFHR), a new mixed-modality dual-reference task in image retrieval where a query consists of a face image specifying identity and a hairstyle reference expressed as either an image or text. Unlike prior retrieval settings, DFHR requires cross-component reasoning between two semantically independent attributes -- identity and hairstyle -- originating from heterogeneous modalities. This formulation demands localized feature disentanglement, cross-modal semantic alignment, and mixed-modality composition within a unified embedding space. We construct DFHR-Bench, the first benchmark for mixed-modality face-hair retrieval, comprising over 180K annotated triplets across dual-image and image-text settings, built via a multi-stage annotation protocol ensuring semantic and identity integrity. We further propose MFHC (Multimodal Face-Hair Combiner), a unified framework that fuses disentangled identity and hairstyle embeddings through token injection and multi-view supervision. DFHR and DFHR-Bench together establish a new paradigm for identity-aware, attribute-controllable visual retrieval across modalities.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Dual Face-Hair Retrieval | DFHR-Bench Bimodal Alignment | Recall@124.82 | 19 | |
| Image-Text Retrieval | DFHR-Bench (Official Set) | Recall@114.7 | 11 | |
| Image+Text Dual Face-Hair Retrieval | DFHR-Bench Image+Text official set | R@114.7 | 11 | |
| Image-to-Image Retrieval | Official Set | Recall@124.72 | 8 | |
| Image+Image Hairstyle Retrieval | DFHR-Bench 1.0 (official set) | R@124.72 | 8 |