SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
About
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Skill execution | SkillsBench | Overall Success Rate (avg@5)43.6 | 26 | |
| Agent task execution | SkillsBench 11 domains | Overall Score32.2 | 16 | |
| Skill-Augmented Agent Performance Evaluation | SkillsBench 86 runnable tasks (frozen-policy transfer) | Valid Task Count76 | 15 | |
| Interactive Agent Task Completion | SocialMaze | Average Reward80.4 | 14 | |
| Interactive Agent Task Completion | ScienceWorld | Average Reward88 | 12 | |
| Downstream task execution | SkillsBench | Reward Mean (%)26.2 | 6 | |
| Downstream task execution | AgentSkillOS | Score78.45 | 6 | |
| Skill-Augmented Agent Performance Evaluation | SkillsBench 20-task balanced panel (held-out evaluation) | Valid Task Count18 | 5 | |
| Offer Letter | SkillsBench | Tokens Used4.10e+4 | 4 | |
| Skill-assisted task execution | SkillsBench 1.0 (test) | Pass@119.5 | 4 |