Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Safety Alignment benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Safety Alignment
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
HarmBench
No Steering
ASR
0
88
4mo ago
Salad Bench
ShaPO-T
MD
0.68
68
4mo ago
HH-RLHF
ShaPO-T
MD Rate
1.09
68
4mo ago
Do-Not-Answer
ShaPO-R
MD
0
52
4mo ago
WildJailbreak
Full FT
Trainable parameters (M)
15,768.31
44
4mo ago
Visual Adversarial Attacks
Vanilla
ASR
43.1
40
3mo ago
JOOD
MoRAS
ASR
0
40
3mo ago
SORRY-Bench
LED-Merging
ASR
10.22
40
2mo ago
PKU-SafeRLHF 30K (IID)
ShaPO-T
WR
89.26
36
4mo ago
AdvBench
SEA
Reward
-0.38
32
4mo ago
Harmful Dataset (test)
Non-Aligned
Harmful Score
81
30
4mo ago
BeaverTails V
SaFeR-ToolKit (+ SFT+GRPO) [3B]
Safety Score
93.37
27
1mo ago
AdvBench
BoN64
Harm Rate
0
25
1mo ago
WildJailbreak
R1 - 8B + UnsafeChain full
Safe@1
77.2
24
3mo ago
Safety Benchmarks (Sorry-bench, StrongREJECT, WildJailbreak, JBB-PAIR, JBB-GCG)
SafeChain
Average Score
42.34
21
2mo ago
XSTest
Yi-VL-6B
Compliance
95.2
21
2mo ago
HEx-PHI
DiaBlo
HEx-PHI Score
98.8
18
2mo ago
HarmBench
SFT
MD Score
95
18
4mo ago
Average (Do-Not-Answer, HarmBench, HH-RLHF, Salad Bench)
ShaPO-T
Aggregate Score
0.59
18
4mo ago
StrongReject
R1 - 7B + UnsafeChain full
Safe@1
58
18
2mo ago
VSA Violence (visual synonym attack prompts)
AEGIS
ASR
0
16
17d ago
RAB Violence (adversarial prompts)
AEGIS
ASR
1
16
17d ago
I2P Violence (explicit prompts)
SDID
ASR
1
16
17d ago
SORRY-Bench
πref
Score
85.7
14
18d ago
PKU-SafeRLHF
PPO
Gold Reward
3.92
14
4mo ago
Showing 25 of 60 rows
25 / page
50 / page
100 / page
1
2
3
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs