Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Code Generation on HumanEval (Model Performance Scores)
Loading...
87.8
HumanEval Score (ID 4)
ANN
71.056
75.403
79.75
84.097
Jun 10, 2025
HumanEval Score (ID 4)
GPT-3.5 Score
GPT-4o-mini Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
HumanEval Score (ID 4)
GPT-3.5 Score
GPT-4o-mini Score
ANN
2025.06
87.8
72.7
90.9
Symbolic
2025.06
85.8
64.5
-
Agents
2025.06
85
59.5
-
Agents w/ AutoPE
2025.06
82.3
63.5
-
DSPy / ToT
2025.06
77.3
66.7
-
GPTs
2025.06
71.7
59.2
-
Feedback
Search any
task
Search any
task