Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Metaphors on Metaphors (Preference accuracy)
Loading...
74.01
Preference Accuracy
DICAI
38.6396
47.8223
57.005
66.1877
Jun 26, 2026
Preference Accuracy
Updated 27d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
DICAI
Judge Model=GPT-4o
2026.06
74.01
ICAI
Judge Model=GPT-4o
2026.06
71.4
CoT-SC
Judge Model=GPT-4o
2026.06
60
ToT
Judge Model=GPT-4o
2026.06
60
Self-Refine
Judge Model=GPT-4o
2026.06
60
AutoRubric
Judge Model=GPT-4o
2026.06
60
CoT
Judge Model=GPT-4o
2026.06
40
Feedback
Search any
task
Search any
task