Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Don't Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance

About

State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples under the same conditioning. Existing methods address this issue via either latent guidance, which has limited effectiveness, or sample selection, which relies on external reward models that incur significant inference-time overhead. In this work, we introduce an efficient, training-free self-guidance mechanism to mitigate diversity collapse without requiring additional reward models. Specifically, we disperse the internal features of the flow model during batch generation with feature self-guidance. Further, to keep the features close to the manifold, we introduce a manifold regularization step that projects these dispersed features back onto the data manifold, ensuring diverse generation without sacrificing alignment with the input conditions. Our method integrates seamlessly as a plug-and-play module into pretrained flow models, adding only a marginal inference cost. Experiments demonstrate significant improvements in diversity while preserving fidelity across several conditional flow models, including multi-step and few-step text-to-image, depth-to-image, and reference image generation.

Pradhaan S Bhat, Rishubh Parihar, Abhijnya Bhat, R. Venkatesh Babu• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Image GenerationGenEval (test)--
250
Personalized Image GenerationDreamBooth
DINO Score57
45
Text-to-Image GenerationGenEval 17 (test)
Latency (s)1.7
10
Text-to-Image GenerationDPGBench (test)
DINO Score0.64
9
Text to ImageGenEval
DINO0.68
8
Text-to-Image GenerationGenEval (test)
Latency (s)0.16
8
Depth-to-Image GenerationCOCO
DINO Score0.52
3
Showing 7 of 7 rows

Other info

Follow for update