Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Code Generation on Python (held-out eval)
Loading...
37.7
Pass@1
Qwen2.5-32B-Instruct (frozen)
11.284
18.142
25
31.858
Jun 17, 2026
Pass@1
Updated 1mo ago
Evaluation Results
Method
Method
Links
Pass@1
Qwen2.5-32B-Instruct (frozen)
Configuration=Qwen2.5-...
2026.06
37.7
Qwen2.5-7B base
Configuration=Qwen2.5-...
2026.06
25
Qwen2.5-7B + CARE 5 campaigns
Configuration=Qwen2.5-...
2026.06
20
Qwen2.5-7B + naive 5 campaigns
Configuration=Qwen2.5-...
2026.06
12.3
Feedback
Search any
task
Search any
task