Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Toward Cybersecurity-Expert Small Language Models

About

Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets. To address this gap, we present CyberPal 2.0, a family of cybersecurity-expert small language models (SLMs) ranging from 4B-20B parameters. To train CyberPal 2.0, we generate an enriched chain-of-thought cybersecurity instruction dataset built with our data enrichment and formatting pipeline, SecKnowledge 2.0, which integrates expert-in-the-loop steering of reasoning formats alongside LLM-driven multi-step grounding, yielding higher-fidelity, task-grounded reasoning traces for security tasks. Across diverse cybersecurity benchmarks, CyberPal 2.0 consistently outperforms its baselines and matches or surpasses various open and closed-source frontier models, while remaining a fraction of their size. On core cyber threat intelligence knowledge tasks, our models outperform almost all tested frontier models, ranking second only to Sec-Gemini v1. On core threat-investigation tasks, such as correlating vulnerabilities and bug tickets with weaknesses, our best 20B-parameter model outperforms GPT-4o, o1, o3-mini, and Sec-Gemini v1, ranking first, while our smallest 4B-parameter model ranks second.

Matan Levi, Daniel Ohayon, Ariel Blobstein, Ravid Sagi, Ian Molloy, Yair Allouche• 2025

Related benchmarks

TaskDatasetResultRank
Security EvaluationSecEval
Accuracy72.86
33
Cybersecurity Knowledge EvaluationCybermetric 2000
Accuracy89.95
26
Cybersecurity Knowledge and ReasoningCybersecurity Evaluation Suite (CTI Bench, SecEval, Cyber Metric 2000, CISSP, Adv. CTI)
SecEval Score69.71
12
Advanced Cyber Threat Intelligence AnalysisAdv. CTI
Accuracy89.58
9
Cyber Threat Relationship PredictionCTI Relationship Prediction
Accuracy92.93
9
Multiple-Choice QuestionsCTI Bench MCQ
Accuracy75.71
9
Professional Certification ExaminationCISSP Exams
Accuracy90.4
9
Relationship Consistency ModelingCTI Bench RCM
Accuracy87.4
9
Threat Detection and MitigationCTI Detect & Mitigate
Accuracy70.95
9
Software Weakness Impact MappingWeakness Impact Mapping
Accuracy71.06
9
Showing 10 of 13 rows

Other info

Follow for update