Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Diagnostic Report Generation on Human-Eval 12-language (120 items)
Loading...
7.41
Human Mean Score
MADE
5.2884
5.8392
6.39
6.9408
Jun 5, 2026
Human Mean Score
Automatic Mean Score
Human Win Rate
Automatic Win Rate
Updated 1mo ago
Evaluation Results
Method
Method
Links
Human Mean Score
Automatic Mean Score
Human Win Rate
Automatic Win Rate
MADE
2026.06
7.41
8.1
87.9
96.2
nanobot
Description=strongest...
2026.06
5.93
5.35
40.6
39.2
cot
Description=strongest...
2026.06
5.37
3.85
21.5
14.6
Feedback
Search any
task
Search any
task