Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Code Generation on HumanEval+ (Accuracy, Follow Rate)
Loading...
77.44
Accuracy
Qwen-2.5-7B-Instruct
57.784
62.887
67.99
73.093
Jun 5, 2026
Accuracy
Follow Rate
Updated 1mo ago
Evaluation Results
Method
Method
Links
Accuracy
Follow Rate
Qwen-2.5-7B-Instruct
Stage=Base
2026.06
77.44
-
Qwen-2.5-7B-Instruct
Stage=Post-trained
2026.06
76.22
-
A.X-4.0-Light
Stage=Base
2026.06
75
-
EXAONE-3.5-7.8B-Instruct
Stage=Base
2026.06
75
-
Kanana-1.5-8B-Instruct
Stage=Base
2026.06
75
-
A.X-4.0-Light
Stage=Post-trained
2026.06
75
-
Kanana-1.5-8B-Instruct
Stage=Post-trained
2026.06
75
-
EXAONE-3.5-7.8B-Instruct
Stage=Post-trained
2026.06
73.78
-
Gemma-3-4B-IT
Stage=Base
2026.06
61.59
-
Gemma-3-4B-IT
Stage=Post-trained
2026.06
61.59
-
Llama-3.1-8B-Instruct
Stage=Post-trained
2026.06
59.15
-
Llama-3.1-8B-Instruct
Stage=Base
2026.06
58.54
-
Feedback
Search any
task
Search any
task