Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Gradient-based Adversarial Attacks against Text Transformers

About

We propose the first general-purpose gradient-based attack against transformer models. Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization. We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks. Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs.

Chuan Guo, Alexandre Sablayrolles, Herv\'e J\'egou, Douwe Kiela• 2021

Related benchmarks

TaskDatasetResultRank
Jailbreak AttackAdvBench
ASR70
283
Jailbreak AttackHarmBench (test)
ASRHB54.17
276
Jailbreak AttackStrongReject (test)
Score56.3
64
Token-forcing loss optimizationRandom targets Held-out (val)
Qwen-2.5-7B Loss11.17
56
Trigger inversionSST2
Success Rate0.0333
44
Trigger inversionSST2
Recall6.6667
44
Adversarial AttackJailbreakBench 50% stratified per-category sample (48 requests)
HB ASR0.00e+0
32
Trigger inversionYahoo high poison rate
Success Rate5
26
Trigger inversionYahoo (test)
Trigger Inversion Success Rate1
26
Trigger inversionYahoo Tell me seriously - 3 tokens high poison rate (test)
Inversion Success Rate0.00e+0
14
Showing 10 of 12 rows

Other info

Follow for update