H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
About
Hairstyle transfer has practical applications such as virtual try-on, yet remains challenging when the source and reference exhibit large head-pose discrepancies. We propose H-Adapter, which improves pose robustness by training with a region-specific loss that disentangles hair and non-hair objectives and thereby induces spatially disentangled cross-attention, from which a source-aligned hair edit mask is derived to guide diffusion-based inpainting. Experiments on pose-agnostic and pose-different subsets demonstrate strong quantitative results, including the best FID, $\mathrm{FID}_{\mathrm{CLIP}}$, and CLIP-I under pose differences, while maintaining competitive non-hair preservation and improving qualitative fidelity to fine-grained reference hairstyle details. Beyond source-conditioned transfer, H-Adapter supports practical extensions including text-to-image generation, auxiliary prompt-based hair color control, and compatibility with an identity-preserving IP-Adapter variant. We also introduce a VLM-as-a-judge protocol and observe consistent gains in hairstyle faithfulness, non-hair preservation, and artifact quality.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Pairwise Preference Evaluation | 40-triplet (val) | HFS91 | 15 | |
| Hairstyle Transfer | CelebA-HQ pose-different | FID12.47 | 9 | |
| Hairstyle Transfer | CelebA-HQ pose-agnostic | FID11.56 | 8 | |
| Hairstyle Transfer | CelebA-HQ | HFS Score3.11 | 6 | |
| Hairstyle Transfer | Human Preference Study | Votes3.19e+3 | 6 | |
| Hairstyle Transfer | Hairstyle Transfer (test) | HFS3.72 | 6 | |
| Hairstyle Transfer | Hairstyle Transfer Evaluation Gemini-2.5-Flash (test) | HFS2.68 | 6 | |
| Hairstyle Transfer | Source-Reference Pairs | Runtime (s)3.05 | 4 |