Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Safety Classification benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Safety Classification
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
AUTALIC (test)
GPT-OSS
F1 Score
37.7
64
1mo ago
HoliSafe-Bench
Qwen3-VL 8B
AUROC
0.783
49
2mo ago
UnsafeBench
Gemma
AUROC
80.5
49
2mo ago
SafeRLHF
Qwen3Guard
F1 Score
0.94
48
3mo ago
WildGuardMix (test)
OSS-Safeguard-20B-High
F1 Score
95.3
47
1mo ago
XSTest
Label and intent reward
F1 Score
95.8
46
17d ago
ToxicChat (test)
D2-TimeAttn
Accuracy
97.3
43
2mo ago
WildGuard (test)
WildGuard 7B
F1 Score
88.8
35
29d ago
ToxicChat
Qwen3Guard
F1 Score
0.81
32
29d ago
AEGIS 2
YuFeng-XGuard-Reason-0.6B
F1 Score
86.2
30
17d ago
Wildguardmix
SingGuard-4B
F1 Score
89.54
29
1mo ago
BeaverTails (test)
Separate Guardrail 7B
AUC
94
24
2mo ago
AEGIS 2.0 (test)
DSA:LST
AUC
94
24
2mo ago
OpenAI-moderation (test)
TPC
Accuracy
74.88
23
2mo ago
HoliSafe-Bench
Llama
ECE
8.4
21
2mo ago
UnsafeBench
Llama
ECE
0.061
21
2mo ago
Pre-Ex-Bench
TRACE
Accuracy
94.01
20
1mo ago
ASSEBench
TRACE
Accuracy
92.04
20
1mo ago
XSTest (test)
LEG large
F1
92.91
20
18d ago
OpenAI Moderation
ShieldGemma 27B
F1 Score
81.4
18
29d ago
Combined Chinese Safety Datasets
CHILLGuard-4B
Overall Average Score
77.93
16
1mo ago
Chinese Response Datasets
Qwen3Guard-0.6B-Strict
Beavertails Score
85.11
16
1mo ago
Chinese Prompt Datasets
PolyGuard-7B
PolyG
86.31
16
1mo ago
MultiJail
CREST-BASE
F1 Score
0.9335
15
2mo ago
HarmBench
IBM Granite Guardian 3.2
Recall
100
14
4mo ago
Showing 25 of 94 rows
25 / page
50 / page
100 / page
1
2
3
4
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs