Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Cybersecurity Knowledge Question Answering on MMLU CSec
Loading...
88
CSec Score
RedSage-8B-Seed
63.7264
70.0282
76.33
82.6318
Jan 29, 2026
Feb 22, 2026
Mar 19, 2026
Apr 13, 2026
May 8, 2026
Jun 2, 2026
Jun 27, 2026
CSec Score
Updated 25d ago
Evaluation Results
Method
Method
Links
CSec Score
RedSage-8B-Seed
evaluation_context=Bas...
2026.01
88
RedSage-8B-Base
evaluation_context=Bas...
2026.01
87
RedSage-8B-CFW
evaluation_context=Bas...
2026.01
86
GPT-5
evaluation_context=Lar...
2026.01
86
LLM-based QA agent with external memory
Model Backend=GPT 5.4...
2026.06
85.34
LLM-based QA agent with external memory
Model Backend=GPT-4o M...
2026.06
85.34
Qwen3-32B
evaluation_context=Lar...
2026.01
84
Llama-3.1-8B
evaluation_context=Bas...
2026.01
83
Qwen3-8B-Base
evaluation_context=Bas...
2026.01
83
Foundation-Sec-8B
evaluation_context=Bas...
2026.01
80
Llama-Primus-Base
evaluation_context=Ins...
2026.01
79
RedSage-8B-DPO
evaluation_context=Ins...
2026.01
79
RedSage-8B-Ins
evaluation_context=Ins...
2026.01
78
LLM-based QA agent with external memory
Model Backend=Gemma2-9...
2026.06
77.59
Llama-Primus-Merged
evaluation_context=Ins...
2026.01
76
Foundation-Sec-8B-Instruct
evaluation_context=Ins...
2026.01
76
Qwen3-8B
evaluation_context=Ins...
2026.01
76
DeepHat-V1-7B
evaluation_context=Ins...
2026.01
74
Llama-3.1-8B-Instruct
evaluation_context=Ins...
2026.01
72
Lily-Cybersecurity-7B-v0.2
evaluation_context=Ins...
2026.01
68
LLM-based QA agent with external memory
Model Backend=Phi3-14B...
2026.06
64.66
Feedback
Search any
task
Search any
task