Our new X account is live! Follow @wizwand_team for updates
Home
/
Benchmarks
Attack Detection on Jailbreak Attacks 105K sample set
Loading...
68
Detection Rate
LogReg (Ours)
27.336
37.893
48.45
59.007
Feb 15, 2026
Detection Rate
Updated 4d ago
Evaluation Results
Method
Method
Links
Detection Rate
LogReg (Ours)
input=raw activations,...
2026.02
68
Llama-as-Judge
acronym=LJ, prompting=...
2026.02
60
PromptGuard 2
acronym=PG
2026.02
48.5
LlamaGuard
acronym=LG
2026.02
28.9
Feedback
Search any
task
Search any
task