Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Code Benchmarks

Benchmarks

Task NameDataset NameSOTA ResultTrend
Code GenerationCode Benchmarks LCBv6 & MBPP+
LCBv6 Score32.8
19
Code ReasoningCode Benchmarks HumanEval MBPP
HumanEval72.29
18
Code GenerationCode Benchmarks (HumanEval, MBPP)
HumanEval Score30.1
17
CodingCode Benchmarks Aggregate
Score37.5
12
Code GenerationCode Benchmarks Total
Avg@826.4
4
Code GenerationCode Benchmarks HumanEval, HumanEval+, MBPP, MBPP+ (test)
HumanEval Score55.5
3
Code GenerationCode Benchmarks (test)
HumanEval34.1
3
Showing 7 of 7 rows