CellularSpecSec-Bench: A Staged Benchmark for Evidence-Grounded Interpretation and Security Reasoning over 3GPP Specifications
About
Cellular networks are critical infrastructure supporting billions of worldwide users and safety- and mission-critical services. Vulnerabilities in cellular networks can therefore cause service disruption, privacy breaches, and broad societal harm, motivating growing efforts to analyze 3GPP specifications that define required device and operator behavior. While large language models (LLMs) have demonstrated the capability for reading technical documents, cellular specifications impose unique challenges: faithful interpretation of normative language, reasoning across cross-referenced clauses, and verifiable conclusions grounded in multimodal evidence such as tables and figures. To address these challenges, we propose CellSpecSec-ARI, a unified Adapt-Retrieve-Integrate framework for systematic understanding and standard-driven security analysis of 3GPP specifications; CellularSpecSec-Bench, a staged benchmark, containing newly constructed high-quality datasets with expert-verified and corrected subsets from prior open-source resources. Together, they establish an accessible and reproducible foundation for quantifying progress in specification understanding and security reasoning in the cellular network security domain.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Vulnerability/Inconsistency Labeling | CellularSpecSec-Bench Stage 3 (Vulnerability / Inconsistency Labeling) | Binary F193.05 | 4 | |
| Abstractive QA | CellularSpecSec-Bench Stage 1 | Score 297 | 2 | |
| Cross-Clause QA (CCQA) | CellularSpecSec-Bench Stage 3 | Score=295.06 | 2 | |
| Evidence and Explanations Correctness | CellularSpecSec-Bench Stage 3 | E&E Correctness88 | 2 | |
| Evidence-Grounded Abstractive QA | CellularSpecSec-Bench Stage 2 | Score (Stage 2)96.8 | 2 | |
| Evidence-Grounded Extractive QA | CellularSpecSec-Bench Stage 2 | Score 20.968 | 2 | |
| Evidence-Grounded MCQA | CellularSpecSec-Bench Stage 2 | Accuracy100 | 2 | |
| Evidence-Grounded Telecommunications QA | TSpec-LLM | Accuracy94 | 2 | |
| Extractive QA | CellularSpecSec-Bench Stage 1 | Stage 1 Score (Component 2)97.75 | 2 | |
| Multiple Choice Question Answering (MCQA) | CellularSpecSec-Bench Stage 1 | Accuracy100 | 2 |