Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning

About

As Multimodal Large Language Models (MLLMs) develop, their potential security issues have become increasingly prominent. Machine Unlearning (MU), as an effective strategy for forgetting specific knowledge in training data, has been widely used in privacy protection. However, MU for safety in MLLM has yet to be fully explored. To address this issue, we propose SAFEERASER, a safety unlearning benchmark for MLLMs, consisting of 3,000 images and 28.8K VQA pairs. We comprehensively evaluate unlearning methods from two perspectives: forget quality and model utility. Our findings show that existing MU methods struggle to maintain model performance while implementing the forget operation and often suffer from over-forgetting. Hence, we introduce Prompt Decouple (PD) Loss to alleviate over-forgetting through decouple prompt during unlearning process. To quantitatively measure over-forgetting mitigated by PD Loss, we propose a new metric called Safe Answer Refusal Rate (SARR). Experimental results demonstrate that combining PD Loss with existing unlearning methods can effectively prevent over-forgetting and achieve a decrease of 79.5% in the SARR metric of LLaVA-7B and LLaVA-13B, while maintaining forget quality and model utility. Our code and dataset will be released upon acceptance. Warning: This paper contains examples of harmful language and images, and reader discretion is recommended.

Junkai Chen, Zhijie Deng, Kening Zheng, Yibo Yan, Shuliang Liu, PeiJun Wu, Peijie Jiang, Jia Liu, Xuming Hu• 2025

Related benchmarks

TaskDatasetResultRank
Visual Question AnsweringVizWiz
Accuracy56.5
1525
Object Hallucination EvaluationPOPE--
1455
Science Question AnsweringScienceQA (SQA)
Accuracy72.2
273
Visual ReasoningGQA
Accuracy62.6
93
Multi-modal UnderstandingMMBench EN
Accuracy68.3
64
Visual Question AnsweringVQA
Accuracy62.3
52
Multimodal Machine Unlearning EvaluationMLLMU-Bench Forget Set
Classification Accuracy51.87
36
Multimodal Machine UnlearningRetain Set
Classification Accuracy48.06
35
Multimodal Machine Unlearning EvaluationMLLMU-Bench Real Celebrity
Class Acc51.8
28
Multimodal Machine Unlearning EvaluationMLLMU-Bench (test)
Classification Accuracy47.86
27
Showing 10 of 16 rows

Other info

Code

Follow for update