Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Error Recognition on WikiMQA
Loading...
44.1
Acc@1
CORRECT
4.9336
15.1018
25.27
35.4382
Sep 28, 2025
Acc@1
Acc@3
Acc@5
Updated 1mo ago
Evaluation Results
Method
Method
Links
Acc@1
Acc@3
Acc@5
CORRECT
Synthesized by=GPT-4o-...
2025.09
44.1
88.2
94.1
CORRECT
Synthesized by=GPT-5-Nano
2025.09
16.9
49.9
69
LLM-as-a-Judge
Synthesized by=GPT-4o-...
2025.09
14.7
55.9
64.7
Baseline
Synthesized by=GPT-5-Nano
2025.09
6.44
35.6
56.1
Feedback
Search any
task
Search any
task