Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MMA-Diffusion: MultiModal Attack on Diffusion Models

About

In recent years, Text-to-Image (T2I) models have seen remarkable advancements, gaining widespread adoption. However, this progress has inadvertently opened avenues for potential misuse, particularly in generating inappropriate or Not-Safe-For-Work (NSFW) content. Our work introduces MMA-Diffusion, a framework that presents a significant and realistic threat to the security of T2I models by effectively circumventing current defensive measures in both open-source models and commercial online services. Unlike previous approaches, MMA-Diffusion leverages both textual and visual modalities to bypass safeguards like prompt filters and post-hoc safety checkers, thus exposing and highlighting the vulnerabilities in existing defense mechanisms.

Yijun Yang, Ruiyuan Gao, Xiaosen Wang, Tsung-Yi Ho, Nan Xu, Qiang Xu• 2023

Related benchmarks

TaskDatasetResultRank
Red TeamingI2P Nudity prompts
Failure Rate (FR)0.86
48
Red Teamingviolence prompts
Failure Rate (FR)6.02
48
Nudity ErasureNudity Erasure
ASR67
48
JailbreakingMHSC
ASR-426.5
44
JailbreakingQ16
ASR-429.5
44
Radiology Report GenerationIU-Xray
ROUGE-L Score32.95
38
JailbreakingUnsafe Prompts
Bypass Success Rate (Text)58
22
Textual Modal AttackLAION-COCO subset, UnsafeDiff, and I2P NSFW prompts (test)
Q16 ASR (Step 4)84.9
15
Jailbreak AttackVBCDE
ASR8
12
Jailbreak AttackUnsafeDiff
Attack Success Rate (ASR)7.3
12
Showing 10 of 36 rows

Other info

Follow for update