Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models

About

Diffusion models show remarkable image generation performance following text prompts, but risk generating sexual contents. Existing approaches, such as prompt filtering, concept removal, and even sexual contents mitigation methods, struggle to defend against adversarial attacks while maintaining benign image quality. In this paper, we propose a novel approach called Distorting Embedding Space (DES), a text encoder-based defense mechanism that effectively tackles these issues through innovative embedding space control. DES transforms unsafe embeddings, extracted from a text encoder using unsafe prompts, toward carefully calculated safe embedding regions to prevent unsafe contents generation, while reproducing the original safe embeddings. DES also neutralizes the ``nudity'' embedding, by aligning it with neutral embedding to enhance robustness against adversarial attacks. As a result, extensive experiments on explicit content mitigation and adaptive attack defense show that DES achieves state-of-the-art (SOTA) defense, with attack success rate (ASR) of 9.47% on FLUX.1, a recent popular model, and 0.52% on the widely adopted Stable Diffusion v1.5. These correspond to ASR reductions of 76.5% and 63.9% compared to previous SOTA methods, EraseAnything and AdvUnlearn, respectively. Furthermore, DES maintains benign image quality, achieving Frechet Inception Distance and CLIP score comparable to those of the original FLUX.1 and Stable Diffusion v1.5.

Jaesin Ahn, Heechul Jung• 2025

Related benchmarks

TaskDatasetResultRank
Compositional Image GenerationGenEval
Overall Score56.7
94
Defense against adversarial promptsMMA (test)
Attack Success Rate0.4
15
Defense against adversarial promptsP4D (test)
Attack Success Rate1.1
15
Defense against adversarial promptsRing-A-Bell (test)
Attack Success Rate0.93
15
Defense against adversarial promptsSneaky (test)
Attack Success Rate0.00e+0
15
Image-to-Image EditingImagenHub Instruction-image pairs
VQA Accuracy84.92
14
Concept ErasureNude Unsafe-1k I2P (test)
CLIP Score31.3
11
Concept ErasureBloody (test)
CLIP Score31.4
11
Concept ErasureVanGogh (test)
CLIP Score31.4
11
Concept ErasurePikachu (test)
CLIP Score31.3
11
Showing 10 of 23 rows

Other info

Follow for update