Gradient-based Adversarial Attacks against Text Transformers
About
We propose the first general-purpose gradient-based attack against transformer models. Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization. We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks. Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs.
Chuan Guo, Alexandre Sablayrolles, Herv\'e J\'egou, Douwe Kiela• 2021
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Jailbreak Attack | AdvBench | ASR70 | 283 | |
| Jailbreak Attack | HarmBench (test) | ASRHB54.17 | 276 | |
| Jailbreak Attack | StrongReject (test) | Score56.3 | 64 | |
| Token-forcing loss optimization | Random targets Held-out (val) | Qwen-2.5-7B Loss11.17 | 56 | |
| Trigger inversion | SST2 | Success Rate0.0333 | 44 | |
| Trigger inversion | SST2 | Recall6.6667 | 44 | |
| Adversarial Attack | JailbreakBench 50% stratified per-category sample (48 requests) | HB ASR0.00e+0 | 32 | |
| Trigger inversion | Yahoo high poison rate | Success Rate5 | 26 | |
| Trigger inversion | Yahoo (test) | Trigger Inversion Success Rate1 | 26 | |
| Trigger inversion | Yahoo Tell me seriously - 3 tokens high poison rate (test) | Inversion Success Rate0.00e+0 | 14 |
Showing 10 of 12 rows