DiverseFlow: Sample-Efficient Diverse Mode Coverage in Flows
About
Many real-world applications of flow-based generative models desire a diverse set of samples that cover multiple modes of the target distribution. However, the predominant approach for obtaining diverse sets is not sample-efficient, as it involves independently obtaining many samples from the source distribution and mapping them through the flow until the desired mode coverage is achieved. As an alternative to repeated sampling, we introduce DiverseFlow: a training-free approach to improve the diversity of flow models. Our key idea is to employ a determinantal point process to induce a coupling between the samples that drives diversity under a fixed sampling budget. In essence, DiverseFlow allows exploration of more variations in a learned flow model with fewer samples. We demonstrate the efficacy of our method for tasks where sample-efficient diversity is desirable, such as text-guided image generation with polysemous words, inverse problems like large-hole inpainting, and class-conditional image synthesis.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Code Generation | HumanEval | -- | 171 | |
| Mathematical Reasoning | GSM8K 200 PROBLEMS | Pass@1683.9 | 65 | |
| Text-to-Image Generation | COCO truck concept 'a photo of a truck' prompt (test) | BRISQUE22.18 | 24 | |
| Class-conditional Image Generation | truck concept | BRISQUE18.73 | 18 | |
| Sampling diversity and quality estimation | Gaussian Mixture | Mode Coverage (Marginal)6.5 | 17 | |
| Text-to-Image Generation | bus concept | BRISQUE23 | 15 | |
| Text-to-Image Generation | bicycle concept (test) | BRISQUE21.19 | 15 | |
| Text-to-Image Generation | "a photo of a apple" prompt apple concept (test) | BRISQUE19.69 | 12 | |
| Text-to-Image Generation | pizza concept a photo of a pizza prompt | BRISQUE20.36 | 12 | |
| In-batch diverse text-to-image generation | MS-COCO prompts | FID39.78 | 6 |