Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Python

Benchmarks

Task NameDataset NameSOTA ResultTrend
Secure Code GenerationPython (test)
Safe Code Rate94.4
66
Graphical API RecommendationPython 13
mAP83.4
56
Speculative DecodingPython
MAT7.69
30
Code SearchPython (test)
Recall@149.4
25
Code Summarization Factual ConsistencyPython
Pearson Correlation (rp)0.497
15
Adversarial Code CompliancePython (Py)
Decoupling Probability99.4
9
Code GenerationPython deduplicated retrieval codebase (test)
EM1,183
9
Code Comment GenerationPython (test)
BLEU21.01
8
Code GenerationPython 10 sequential 20-step campaigns (end-of-chain)
End pass@125.9
5
Code SearchPython
MRR85.63
5
Code GenerationPython (held-out eval)
Pass@137.7
4
Package-level evaluationPython Full (val)
Precision92.4
4
RRT-Connect PlanningPython 2D
Wall-Clock Time (ms)4.4
2
Showing 13 of 13 rows