Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Benchmark Score Prediction on model-benchmark score matrix (held-out cells)

5.86MedAPE

GPT-5.5

5.77886.32696.8757.4231Jun 22, 2026
Updated 1mo ago

Evaluation Results

MethodLinks
2026.06
5.863.5
2026.06
7.774.63
2026.06
7.894.7