Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

APIBench

Benchmarks

Task NameDataset NameSOTA ResultTrend
API UsageAPIBench HF (test)
Accuracy83.2
12
Tool InvocationAPIBench
HuggingFace AST Accuracy79.2
8
Tool UseAPIBench OOD - TensorHub
Hallucination Rate2.04
6
Tool UseAPIBench OOD TorchHub
Hallucination Rate5.91
6
Tool UseAPIBench OOD HuggingFace
Hallucination Rate0.0642
6
Tool SelectionAPIBench (test)
Recall@130.64
4
API Question AnsweringAPIBench TensorFlow Hub (test)
Accuracy88.91
4
API Question AnsweringAPIBench Torch Hub (test)
Accuracy93.55
4
API Question AnsweringAPIBench Hugging Face (test)
Accuracy77.21
4
Showing 9 of 9 rows