Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation

About

Diffusion-based models have gained significant popularity for text-to-image generation due to their exceptional image-generation capabilities. A risk with these models is the potential generation of inappropriate content, such as biased or harmful images. However, the underlying reasons for generating such undesired content from the perspective of the diffusion model's internal representation remain unclear. Previous work interprets vectors in an interpretable latent space of diffusion models as semantic concepts. However, existing approaches cannot discover directions for arbitrary concepts, such as those related to inappropriate concepts. In this work, we propose a novel self-supervised approach to find interpretable latent directions for a given concept. With the discovered vectors, we further propose a simple approach to mitigate inappropriate generation. Extensive experiments have been conducted to verify the effectiveness of our mitigation approach, namely, for fair generation, safe generation, and responsible text-enhancing generation. Project page: \url{https://interpretdiffusion.github.io}.

Hang Li, Chengzhi Shen, Philip Torr, Volker Tresp, Jindong Gu• 2023

Related benchmarks

TaskDatasetResultRank
Text-to-Image GenerationMS-COCO
FID72.56
193
Text-to-Image GenerationCOCO 30k
FID15.98
77
Concept UnlearningUnlearnDiffAtk
UnlearnDiffAtk0.697
36
Safe Text-to-Image GenerationMMA-Diffusion
Automatic Safety Rate90.7
33
Text-to-Image GenerationRAB
ASR0.89
21
Text-to-Image GenerationVSA
ASR73
21
Text-to-Image GenerationI2P
ASR33
21
Text-to-Image GenerationMMA
ASR66
21
Fair GenerationWinoBias Gender-Pro extended
Deviation Ratio0.07
20
Fair GenerationWinoBias Race (standard)
Deviation Ratio0.04
20
Showing 10 of 42 rows

Other info

Code

Follow for update