EVA: Editing for Versatile Alignment against Jailbreaks

About

Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks, where adversaries exploit textual or visual triggers to bypass safety guardrails. Recent defenses typically rely on safety fine-tuning or external filters to reduce the model's likelihood of producing harmful content. While effective to some extent, these methods often incur significant computational overheads and suffer from the safety utility trade-off, degrading the model's performance on benign tasks. To address these challenges, we propose EVA (Editing for Versatile Alignment against Jailbreaks), a novel framework that pioneers the application of direct model editing for safety alignment. EVA reframes safety alignment as a precise knowledge correction task. Instead of retraining massive parameters, EVA identifies and surgically edits specific neurons responsible for the model's susceptibility to harmful instructions, while leaving the vast majority of the model unchanged. By localizing the updates, EVA effectively neutralizes harmful behaviors without compromising the model's general reasoning capabilities. Extensive experiments demonstrate that EVA outperforms baselines in mitigating jailbreaks across both LLMs and VLMs, offering a precise and efficient solution for post-deployment safety alignment.

Yi Wang, Hongye Qiu, Yue Xu, Sibei Yang, Zhan Qin, Minlie Huang, Wenjie Wang• 2026

Related benchmarks

Task	Dataset	Result
Natural Language Inference	RTE	Accuracy83.3	590
Multimodal Evaluation	MMStar	Accuracy67.6	177
Named Entity Recognition	CoNLL 03	F1 Score0.51	140
Jailbreak Defense	AdvBench	ASR (PAIR)0.00e+0	115
Multi-turn conversation	MT-Bench	Average Score8.48	107
Jailbreak Defense	HarmBench	PAIR ASR0.00e+0	91
Reasoning	GSM8K	Accuracy (GSM8K)98.8	55
Multimodal Evaluation	MM-Vet v2	Score69.3	46
Dialogue Reasoning	MuTual	Accuracy79.8	38
Multimodal Evaluation	MMMU	Score57.6	36

Showing 10 of 16 rows

Other info

Follow for update

@wizwand_team Discord