Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Harmful Content Detection on ToxicChat
Loading...
92.6
Safe Rate
Llama Guard 3
33.008
48.479
63.95
79.421
Jul 16, 2025
Safe Rate
Unsafe Rate
Updated 18d ago
Evaluation Results
Method
Method
Links
Safe Rate
Unsafe Rate
Llama Guard 3
2025.07
92.6
47.2
Latent Guard-Llama3
Backbone=Llama3
2025.07
83.5
31.7
Latent Guard-Qwen2
Backbone=Qwen2
2025.07
80.1
34
Latent Guard-Llama2
Backbone=Llama2
2025.07
35.3
72.7
Feedback
Search any
task
Search any
task