Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

About

Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization. This leaves a blind spot for visual synonym attacks (VSA), a jailbreak where benign-looking prompts elicit prohibited imagery through implicit visual associations. As a result, current defenses face a safety-utility dilemma: they may either under-mitigate VSA threats or over-suppress visually similar benign concepts. The core challenge is that VSA hides the unsafe target at the textual surface while revealing it through generation-time visual-semantic convergence. In this work, we therefore shift from static suppression of pre-specified unsafe concepts to dynamic tracing of how unsafe semantics emerge during generation. Our mechanistic analysis shows that VSA and explicit unsafe prompts converge through sparse semantic-injecting attention heads, which serve as inference-time bottlenecks for prohibited visual semantics. Based on this insight, we propose AEGIS (Adaptive Evasion Guard via Identification and Steering), an inference-time defense that applies similarity-aware repulsion only at the identified vulnerable heads. Evaluated against 16 baselines, AEGIS improves both safety and utility. On SD 1.4, it reduces ASR to $\mathbf{0.00}/\mathbf{0.03}$ for in-domain violence/nudity VSA and achieves ASRs $\le \mathbf{0.09}$ on out-of-domain explicit and adversarial attacks. It preserves benign fidelity, avoids suppressing hard-negative concepts, and transfers to SD 2.1 and FLUX.1 after re-identifying the critical heads for each backbone.

Yuanmin Huang, Zhenfei Zhang, Mi Zhang, Geng Hong, Qinqin He, Jialing Tao, Hui Xue, Min Yang• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Image GenerationMS-COCO
FID68.25
193
Text-to-Image GenerationRAB
ASR0.17
21
Text-to-Image GenerationI2P
ASR2
21
Text-to-Image GenerationMMA
ASR1
21
Text-to-Image GenerationVSA
ASR3
21
Semantic PreservationHard Negative Prompts (Nudity)
CLIP Score35.32
19
Safety AlignmentRAB Violence (adversarial prompts)
ASR1
16
Safety AlignmentVSA Violence (visual synonym attack prompts)
ASR0.00e+0
16
Safety AlignmentI2P Violence (explicit prompts)
ASR1
16
Utility PreservationStandard Benign Prompts
CLIP Score31.09
16
Showing 10 of 14 rows

Other info

Follow for update