Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MCP-Bench

Benchmarks

Task NameDataset NameSOTA ResultTrend
Tool-augmented agent executionMCP-Bench
Task Fulfillment53.5
32
Agentic Tool UseMCP-Bench 3-server 1.0
Overall Score3.95
20
Agentic Tool UseMCP-Bench Single 1.0
Overall Score3.54
20
Agent Performance EvaluationMCP-Bench
Task Fulfillment46.8
7
End-to-End Question AnsweringMCP-Bench
Accuracy (Human)87.5
4
Showing 5 of 5 rows