Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

About

With the significant advancement of Large Vision-Language Models (VLMs), concerns about their potential misuse and abuse have grown rapidly. Previous studies have highlighted VLMs' vulnerability to jailbreak attacks, where carefully crafted inputs can lead the model to produce content that violates ethical and legal standards. However, existing methods struggle against state-of-the-art VLMs like GPT-4o, due to the over-exposure of harmful content and lack of stealthy malicious guidance. In this work, we propose a novel jailbreak attack framework: Multi-Modal Linkage (MML) Attack. Drawing inspiration from cryptography, MML utilizes an encryption-decryption process across text and image modalities to mitigate over-exposure of malicious information. To align the model's output with malicious intent covertly, MML employs a technique called "evil alignment", framing the attack within a video game production scenario. Comprehensive experiments demonstrate MML's effectiveness. Specifically, MML jailbreaks GPT-4o with attack success rates of 97.80% on SafeBench, 98.81% on MM-SafeBench and 99.07% on HADES-Dataset. Our code is available at https://github.com/wangyu-ovo/MML.

Yu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang, Tianxing He• 2024

Related benchmarks

Task	Dataset	Result
Jailbreak Attack	HarmBench	--	557
Jailbreak Attack	SafeBench	ASR22.4	245
Jailbreak Attack	JailbreakBench	ASR0.00e+0	242
Jailbreak Attack	Malicious goals dataset (test)	ASR0.00e+0	99
Jailbreak Safety Evaluation	MM-Safety Bench (test)	Average ASR14.92	56
Multimodal Jailbreak Attack	MM-SafetyBench (full)	ASR95.42	40
Jailbreak Attack	Claude 3.5	ASR60.4	24
Multimodal Jailbreaking	HADES-Dataset	ASR (%)99.07	20
Jailbreaking	GPT-4o	ASR0.978	19
Safety Evaluation	SafeBench	Overall Safety Score88	19

Showing 10 of 25 rows

Other info

Code

Follow for update

@wizwand_team Discord