Our new X account is live! Follow @wizwand_team for updates
Home
/
Benchmarks
Task Performance Evaluation on Cybersecurity
Loading...
86.2
Average Score
CASTER
81
82.35
83.7
85.05
Jan 27, 2026
Average Score
Quality Gain
Updated 4d ago
Evaluation Results
Method
Method
Links
Average Score
Quality Gain
CASTER
Strategy=CASTER
2026.01
86.2
-
Force Strong
Strategy=Force Strong
2026.01
85.5
-
Force Weak
Strategy=Force Weak
2026.01
83.5
-
CASTER
2026.01
82
0.8
FrugalGPT (Cascade)
2026.01
81.2
-
Feedback
Search any
task
Search any
task