Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DevEval

Benchmarks

Task NameDataset NameSOTA ResultTrend
Developer Knowledge EvaluationDevEval
Win Rate61
7
Docstring EvaluationDevEval 183 human-written docstrings
Score4.938
5
Agentic CodingDevEval
Solve Rate94.8
4
Repository-level code generationDevEval
Inference Time442
4
Terminal-related CLI agent taskDevEval
Accuracy39.74
2
Showing 5 of 5 rows