Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Preference Reconstruction on Experiment Design
Loading...
81.77
Preference Accuracy
DICAI
54.4076
61.5113
68.615
75.7187
Jun 26, 2026
Preference Accuracy
Updated 26d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
DICAI
Judge LLM=GPT-5
2026.06
81.77
DICAI
Judge Model=GPT-4o
2026.06
79.93
ICAI
Judge Model=GPT-4o
2026.06
73.8
ICAI
Judge LLM=GPT-5
2026.06
73.4
CoT
Judge LLM=GPT-5
2026.06
66.22
CoT
Judge Model=GPT-4o
2026.06
65.66
CoT-SC
Judge LLM=GPT-5
2026.06
64.96
ToT
Judge LLM=GPT-5
2026.06
64.87
Self-Refine
Judge Model=GPT-4o
2026.06
64.53
CoT-SC
Judge Model=GPT-4o
2026.06
63.31
Self-Refine
Judge LLM=GPT-5
2026.06
62.52
ToT
Judge Model=GPT-4o
2026.06
62.41
AutoRubric
Judge LLM=GPT-5
2026.06
57.12
AutoRubric
Judge Model=GPT-4o
2026.06
55.46
Feedback
Search any
task
Search any
task