Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Language Modeling on MLGym Language Modeling
Loading...
3.5
Loss
Human Best
3.45308
3.76979
4.0865
4.40321
Jun 20, 2026
Loss
Updated 1mo ago
Evaluation Results
Method
Method
Links
Loss
Human Best
Method Type=Human Base...
2026.06
3.5
ARTS*
Base Model=Qwen 4B, Te...
2026.06
3.518
ARTS
Base Model=o3
2026.06
3.827
Linear
Method Type=Prior Works
2026.06
3.986
MLEvolve
Method Type=Prior Works
2026.06
4.015
ARTS
Base Model=Qwen 4B
2026.06
4.34
AIRA
Method Type=Prior Works
2026.06
4.673
Feedback
Search any
task
Search any
task