Toward Cybersecurity-Expert Small Language Models
About
Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets. To address this gap, we present CyberPal 2.0, a family of cybersecurity-expert small language models (SLMs) ranging from 4B-20B parameters. To train CyberPal 2.0, we generate an enriched chain-of-thought cybersecurity instruction dataset built with our data enrichment and formatting pipeline, SecKnowledge 2.0, which integrates expert-in-the-loop steering of reasoning formats alongside LLM-driven multi-step grounding, yielding higher-fidelity, task-grounded reasoning traces for security tasks. Across diverse cybersecurity benchmarks, CyberPal 2.0 consistently outperforms its baselines and matches or surpasses various open and closed-source frontier models, while remaining a fraction of their size. On core cyber threat intelligence knowledge tasks, our models outperform almost all tested frontier models, ranking second only to Sec-Gemini v1. On core threat-investigation tasks, such as correlating vulnerabilities and bug tickets with weaknesses, our best 20B-parameter model outperforms GPT-4o, o1, o3-mini, and Sec-Gemini v1, ranking first, while our smallest 4B-parameter model ranks second.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Security Evaluation | SecEval | Accuracy72.86 | 33 | |
| Cybersecurity Knowledge Evaluation | Cybermetric 2000 | Accuracy89.95 | 26 | |
| Cybersecurity Knowledge and Reasoning | Cybersecurity Evaluation Suite (CTI Bench, SecEval, Cyber Metric 2000, CISSP, Adv. CTI) | SecEval Score69.71 | 12 | |
| Advanced Cyber Threat Intelligence Analysis | Adv. CTI | Accuracy89.58 | 9 | |
| Cyber Threat Relationship Prediction | CTI Relationship Prediction | Accuracy92.93 | 9 | |
| Multiple-Choice Questions | CTI Bench MCQ | Accuracy75.71 | 9 | |
| Professional Certification Examination | CISSP Exams | Accuracy90.4 | 9 | |
| Relationship Consistency Modeling | CTI Bench RCM | Accuracy87.4 | 9 | |
| Threat Detection and Mitigation | CTI Detect & Mitigate | Accuracy70.95 | 9 | |
| Software Weakness Impact Mapping | Weakness Impact Mapping | Accuracy71.06 | 9 |