Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Jailbreak Attack Evaluation benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Jailbreak Attack Evaluation
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
S-Eval Aattack
Original
Attack Success Rate (ASR)
92
72
4mo ago
SafeBench 100 sampled harmful queries
StructBreak
ASR
97
48
2mo ago
TRIDENT CORE
TRIDENT-CORE
HPR
7
38
4mo ago
DeepIn
ICD
ASR
68
20
18d ago
GPTFuzz
ICD
ASR
92
20
18d ago
HarmBench (400 random samples)
Llama-2-7B-Chat (Original)
ASR
0
18
4mo ago
Malicious Intent Prompts
Timed-Release
ASR
100
16
1mo ago
SafetyBench MCV
Qwen2.5-VL-32B
ASR (1-Clip)
79.79
16
1mo ago
500 randomly sampled prompts (test)
GCG
Similarity Score
0.81
16
2mo ago
ReNe
Intention Analysis
ASR
0
14
18d ago
JBB sampled harmful behaviors
M
PAIR Success Rate
100
12
2mo ago
StealthGraph SG-Implicit
Grok 3 Mini
ASR
91
12
4mo ago
AdvBench
Amnesia
ASR Success Rate
86.3
9
2mo ago
Paired Prompts Held-out (test)
PAIR
Similarity
0.78
8
2mo ago
TRIDENT-EDGE
TRIDENT-EDGE
HPR
5
7
4mo ago
Five Safety Benchmarks AdvBench, HarmBench, HarmfulQ, JBBench, StrongReject
QwQ
ASR
7.69
6
3mo ago
StealthGraph SG-Origin
Mixtral 8×7B
ASR
39.5
6
4mo ago
HarmfulQA
DeepSeek V3.1
ASR
16
6
4mo ago
Do-Not-Answer
Gemini 2.5 Flash
ASR
2.5
6
4mo ago
FigStep Average
SafeThink
Average ASR
0.053
5
4mo ago
AdvBench-X Swahili multilingual (test)
Trajectory-Level Safety Alignment
ASR (OpenAI Moderation)
1.34
3
1mo ago
AdvBench-X Korean multilingual (test)
Trajectory-Level Safety Alignment
ASR (OpenAI Moderation)
5.19
3
1mo ago
POLARIS
POLARIS
Attack Success Count
520
2
2mo ago
Curiosity
POLARIS
Successful Attack Count
5
2
2mo ago
SOS
POLARIS
Successful Attack Count
878
2
2mo ago
Showing 25 of 28 rows
25 / page
50 / page
100 / page
1
2
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs