Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Distraction is All You Need for Multimodal Large Language Model Jailbreaking

About

Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with text can introduce vulnerabilities, which may be exploited to bypass safety mechanisms. To address this, we analyze the relationship between image content and task and find that the complexity of subimages, rather than their content, is key. Building on this insight, we propose the Distraction Hypothesis, followed by a novel framework called Contrasting Subimage Distraction Jailbreaking (CS-DJ), to achieve jailbreaking by disrupting MLLMs alignment through multi-level distraction strategies. CS-DJ consists of two components: structured distraction, achieved through query decomposition that induces a distributional shift by fragmenting harmful prompts into sub-queries, and visual-enhanced distraction, realized by constructing contrasting subimages to disrupt the interactions among visual elements within the model. This dual strategy disperses the model's attention, reducing its ability to detect and mitigate harmful content. Extensive experiments across five representative scenarios and four popular closed-source MLLMs, including GPT-4o-mini, GPT-4o, GPT-4V, and Gemini-1.5-Flash, demonstrate that CS-DJ achieves average success rates of 52.40% for the attack success rate and 74.10% for the ensemble attack success rate. These results reveal the potential of distraction-based approaches to exploit and bypass MLLMs' defenses, offering new insights for attack strategies.

Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Xin Lin, Tao Li, Kanghua Mo, Changyu Dong• 2025

Related benchmarks

TaskDatasetResultRank
Jailbreak AttackHarmBench (test)
ASRHB56.8
276
Jailbreak AttackSafeBench
ASR10.2
245
Jailbreak AttackHADES
Attack Success Rate86.59
132
Jailbreak AttackSafeBench Tiny
ASR56
94
Jailbreaking AttackMM-SafetyBench
Attack Success Rate (ASR)75.06
82
Multimodal Jailbreak AttackHarmBench
ASR9.5
62
JailbreakHarmBench
Toxicity Score1.28
50
Jailbreak AttackClaude 3.5
ASR31.25
24
Jailbreak AttackHADES
Success Rate (Animal)66.67
23
Jailbreak attacksMM-Safety
Safety Rate54
22
Showing 10 of 32 rows

Other info

Follow for update