Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning

About

Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks. More concerningly, prevailing safety tuning strategies tend to provide only superficial safety protection, as they fall short of completely eliminating the backdoor effects. In this work, we present a novel formulation of backdoor learning and unlearning as a sequential, three-stage process from a continual learning perspective. Within this framework, we formally define complete backdoor unlearning and further derive the necessary conditions for achieving it based on the mechanism of catastrophic forgetting. Guided by these insights, we propose Blind Inversion-Backdoor Adversarial Unlearning (BI-BAU), which formulates the generation of adversarial examples satisfying the unlearning conditions as a blind inversion problem. We solve this by integrating the bi-level optimization process of adversarial training into an Expectation-Maximization (EM) algorithm framework to optimize the maximum a posteriori (MAP) objective. Furthermore, BI-BAU is extended to untargeted adversarial scenarios with unknown target classes, as well as to multi-modal contrastive learning tasks, enhancing its applicability to real-world deployment scenarios where pre-trained models may be compromised. Extensive experiments demonstrate that our method exhibits general applicability across a wide spectrum of backdoor attacks and can effectively and thoroughly eliminate the backdoor effects from a backdoor model.

Zhenqian Zhu, Yamin Hu, Yujiang Liu, Luping Wei, Wenbo Hou, Bin Li, Haodong Li, Wenjian Luo• 2026

Related benchmarks

TaskDatasetResultRank
Backdoor DefenseCIFAR10 (test)
ASR2.3
333
Backdoor DefenseGTSRB (test)
ASR0.00e+0
138
Backdoor DefenseGTSRB
CA96.32
118
Backdoor DefenseTiny ImageNet 200 (test)
BadNets CA44.6
11
Backdoor Defense RobustnessTiny-ImageNet-200
BadNets CA41.14
9
Backdoor Defense EvaluationCIFAR-10
CA (BadNets)83.81
9
Backdoor DefenseCIFAR-10 (test)
BEC (BadNets)15.01
7
Showing 7 of 7 rows

Other info

Follow for update