Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Preference Reconstruction on Metaphors
Loading...
75.22
Preference Accuracy
DICAI
59.3912
63.5006
67.61
71.7194
Jun 26, 2026
Preference Accuracy
Updated 27d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
DICAI
Judge LLM=GPT-5
2026.06
75.22
ICAI
Judge LLM=GPT-5
2026.06
74.4
AutoRubric
Judge LLM=GPT-5
2026.06
62.33
CoT
Judge LLM=GPT-5
2026.06
60
CoT-SC
Judge LLM=GPT-5
2026.06
60
ToT
Judge LLM=GPT-5
2026.06
60
Self-Refine
Judge LLM=GPT-5
2026.06
60
Feedback
Search any
task
Search any
task