Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Constrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense

About

We develop a novel optimization method for NLPbackdoor inversion. We leverage a dynamically reducing temperature coefficient in the softmax function to provide changing loss landscapes to the optimizer such that the process gradually focuses on the ground truth trigger, which is denoted as a one-hot value in a convex hull. Our method also features a temperature rollback mechanism to step away from local optimals, exploiting the observation that local optimals can be easily deter-mined in NLP trigger inversion (while not in general optimization). We evaluate the technique on over 1600 models (with roughly half of them having injected backdoors) on 3 prevailing NLP tasks, with 4 different backdoor attacks and 7 architectures. Our results show that the technique is able to effectively and efficiently detect and remove backdoors, outperforming 4 baseline methods.

Guangyu Shen, Yingqi Liu, Guanhong Tao, Qiuling Xu, Zhuo Zhang, Shengwei An, Shiqing Ma, Xiangyu Zhang• 2022

Related benchmarks

TaskDatasetResultRank
Backdoor DetectionSST2
TPR5
56
Backdoor Trigger DetectionSST-2
Recall7
48
Trigger inversionSST2
Success Rate0.0667
44
Trigger inversionSST2
Recall10
44
Trigger inversionYahoo (test)
Trigger Inversion Success Rate5
26
Trigger inversionYahoo high poison rate
Success Rate3
26
Backdoor DetectionYahoo
Tell me seriously2
24
Trigger inversionYahoo
Trigger Inversion Success Rate3
22
Backdoor DetectionSST2 high poison rate (test)
Clean Performance Score8
18
Trigger inversionYahoo Tell me seriously - 3 tokens high poison rate (test)
Inversion Success Rate0.0333
14
Showing 10 of 14 rows

Other info

Follow for update