Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization

About

Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objective, or depend on large-scale and costly RL-trained generators. Motivated by these limitations, we propose JANUS , a lightweight framework that formulates jailbreak as optimizing a structured prompt distribution under a black-box, end-to-end reward from the T2I system and its safety filters. JANUS replaces a high-capacity generator with a low-dimensional mixing policy over two semantically anchored prompt distributions, enabling efficient exploration while preserving the target semantics. On modern T2I models, we outperform state-of-the-art jailbreak methods, improving ASR-8 from 25.30% to 43.15% on Stable Diffusion 3.5 Large Turbo with consistently higher CLIP and NSFW scores. JANUS succeeds across both open-source and commercial models. These findings expose structural weaknesses in current T2I safety pipelines and motivate stronger, distribution-aware defenses. Warning: This paper contains model outputs that may be offensive.

Haolun Zheng, Yu He, Tailun Chen, Shuo Shao, Zhixuan Chu, Hongbin Zhou, Lan Tao, Zhan Qin, Kui Ren• 2026

Related benchmarks

TaskDatasetResultRank
Jailbreak AttackSDXL--
11
Jailbreak AttackSD LT 3.5
TASR (%)94.25
6
Jailbreak AttackDALL-E 3
TASR (%)12.98
6
Jailbreak AttackMidjourney
TASR40.7
6
Jailbreaking Text-to-ImageCivitai NSFW on SD3.5LT
Runtime (s)87.57
6
Showing 5 of 5 rows

Other info

Follow for update