Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Error Recognition on MMLU-Pro
Loading...
69.1
Acc@1
CORRECT
49.236
54.393
59.55
64.707
Sep 28, 2025
Acc@1
Acc@3
Acc@5
Updated 1mo ago
Evaluation Results
Method
Method
Links
Acc@1
Acc@3
Acc@5
CORRECT
Synthesized by=GPT-4o-...
2025.09
69.1
88.2
94.1
CORRECT
Synthesized by=GPT-5-Nano
2025.09
69.1
88.2
94.1
LLM-as-a-Judge
Synthesized by=GPT-4o-...
2025.09
58.3
62.5
66.7
Baseline
Synthesized by=GPT-5-Nano
2025.09
50
64.7
70.59
Feedback
Search any
task
Search any
task