Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

HumanEvalFix

Benchmarks

Task NameDataset NameSOTA ResultTrend
Code RepairHumanEvalFix (test)
Success Rate (Python)59.1
19
Bug-fixingHumanEvalFix held-out 25% (eval)
Utility Difference (Δπ)5.2
3
Showing 2 of 2 rows