Constrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense
About
We develop a novel optimization method for NLPbackdoor inversion. We leverage a dynamically reducing temperature coefficient in the softmax function to provide changing loss landscapes to the optimizer such that the process gradually focuses on the ground truth trigger, which is denoted as a one-hot value in a convex hull. Our method also features a temperature rollback mechanism to step away from local optimals, exploiting the observation that local optimals can be easily deter-mined in NLP trigger inversion (while not in general optimization). We evaluate the technique on over 1600 models (with roughly half of them having injected backdoors) on 3 prevailing NLP tasks, with 4 different backdoor attacks and 7 architectures. Our results show that the technique is able to effectively and efficiently detect and remove backdoors, outperforming 4 baseline methods.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Backdoor Detection | SST2 | TPR5 | 56 | |
| Backdoor Trigger Detection | SST-2 | Recall7 | 48 | |
| Trigger inversion | SST2 | Success Rate0.0667 | 44 | |
| Trigger inversion | SST2 | Recall10 | 44 | |
| Trigger inversion | Yahoo (test) | Trigger Inversion Success Rate5 | 26 | |
| Trigger inversion | Yahoo high poison rate | Success Rate3 | 26 | |
| Backdoor Detection | Yahoo | Tell me seriously2 | 24 | |
| Trigger inversion | Yahoo | Trigger Inversion Success Rate3 | 22 | |
| Backdoor Detection | SST2 high poison rate (test) | Clean Performance Score8 | 18 | |
| Trigger inversion | Yahoo Tell me seriously - 3 tokens high poison rate (test) | Inversion Success Rate0.0333 | 14 |