Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Short Stories on MuCE
Loading...
0.7822
Preference Accuracy
DICAI
0.666448
0.696499
0.72655
0.756601
Jun 26, 2026
Preference Accuracy
Updated 27d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
DICAI
Judge Model=GPT-4o
2026.06
0.7822
CoT
Judge Model=GPT-4o
2026.06
0.7179
ICAI
Judge Model=GPT-4o
2026.06
0.714
AutoRubric
Judge Model=GPT-4o
2026.06
0.6951
ToT
Judge Model=GPT-4o
2026.06
0.6875
Self-Refine
Judge Model=GPT-4o
2026.06
0.675
CoT-SC
Judge Model=GPT-4o
2026.06
0.6709
Feedback
Search any
task
Search any
task