Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Hypothesis Generation
Loading...
77.22
Preference Accuracy
DICAI
51.5736
58.2318
64.89
71.5482
Jun 26, 2026
Preference Accuracy
Updated 27d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
DICAI
Judge Model=GPT-4o
2026.06
77.22
ICAI
Judge Model=GPT-4o
2026.06
64.6
CoT
Judge Model=GPT-4o
2026.06
56.1
ToT
Judge Model=GPT-4o
2026.06
55.42
CoT-SC
Judge Model=GPT-4o
2026.06
54.96
AutoRubric
Judge Model=GPT-4o
2026.06
54.92
Self-Refine
Judge Model=GPT-4o
2026.06
52.56
Feedback
Search any
task
Search any
task