Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Jailbreak Detection benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Jailbreak Detection
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
AdvBench Llama2-7B
GradSafe
AUROC
97.9
88
1mo ago
Average of six attacks
GradSafe
Avg Success Rate
0
38
3mo ago
Jailbreak data (70/30 stratified)
LLMScan
AUC
100
32
2mo ago
Zulu
IMAG
Accuracy
92
30
4mo ago
Base64
IMAG
Accuracy
100
30
4mo ago
DrAttack
GradSafe
Accuracy
99
30
4mo ago
PAIR
IMAG
Accuracy
98
30
4mo ago
AutoDAN
IMAG
Accuracy
99
30
4mo ago
GCG
IMAG
Accuracy
99
30
4mo ago
GoalFrameBench
FrameShield-Crit
Accuracy
94
24
4mo ago
MM-SafetyBench
Mahal-OOD
AUROC
99.18
23
1mo ago
AdvBench Vicuna-7B
MTK
AUROC
95.7
16
1mo ago
AdvBench Mistral-7B
MTK
AUROC
99
16
1mo ago
AdvBench Llama3-8B
GradSafe
AUROC
0.967
16
1mo ago
GoalFrameBench (seed prompts)
FrameShield-Last
Accuracy
97
16
4mo ago
JailbreaksOverTime (test)
MLJailDe
Throughput (items/s)
38.06
15
1mo ago
GCG attacks, Databricks Dolly 15K, and OR-Bench PMPs (test)
SelfDefend
True Positive Rate (TPR)
99
15
1mo ago
DrAttack
SelfDefend (Intent)
ASR
3
15
3mo ago
AutoDAN
GradientCuff
Attack Success Rate (ASR)
0
15
3mo ago
GCG
LLama-3-8B-Instruct (No Defense)
ASR
13
15
3mo ago
ChatGPT Jailbreak Prompts
Llama Prompt Guard 2
Recall
100
15
4mo ago
Wildjailbreak
Apriel Guard
F1 Score
96
15
4mo ago
Jailbreak V28K
MTK
AUROC
96.4
14
1mo ago
OKT
Llama-Prompt-Guard-2
Correlation Score
1
13
4mo ago
SB
gpt-4o-mini
COR
98.33
13
4mo ago
Showing 25 of 94 rows
25 / page
50 / page
100 / page
1
2
3
4
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs