Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Preference Reconstruction on Alternate Uses of Objects
Loading...
78.61
Preference Accuracy
CoT
57.9036
63.2793
68.655
74.0307
Jun 26, 2026
Preference Accuracy
Updated 26d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
CoT
Judge Model=GPT-4o
2026.06
78.61
ToT
Judge LLM=GPT-5
2026.06
77.72
ToT
Judge Model=GPT-4o
2026.06
76.46
Self-Refine
Judge LLM=GPT-5
2026.06
75.35
CoT
Judge LLM=GPT-5
2026.06
74.94
DICAI
Judge Model=GPT-4o
2026.06
74.23
Self-Refine
Judge Model=GPT-4o
2026.06
73.78
DICAI
Judge LLM=GPT-5
2026.06
73.22
CoT-SC
Judge LLM=GPT-5
2026.06
70.77
CoT-SC
Judge Model=GPT-4o
2026.06
70.3
ICAI
Judge LLM=GPT-5
2026.06
69.4
ICAI
Judge Model=GPT-4o
2026.06
66.4
AutoRubric
Judge LLM=GPT-5
2026.06
59.22
AutoRubric
Judge Model=GPT-4o
2026.06
58.7
Feedback
Search any
task
Search any
task