Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Preference Reconstruction on LiTBench Long Stories
Loading...
71.9
Preference Accuracy
Self-Refine
58.7128
62.1364
65.56
68.9836
Jun 26, 2026
Preference Accuracy
Updated 27d ago
Evaluation Results
Method
Method
Links
Preference Accuracy
Self-Refine
Judge Model=GPT-4o
2026.06
71.9
ToT
Judge LLM=GPT-5
2026.06
71.29
DICAI
Judge LLM=GPT-5
2026.06
71.2
CoT
Judge LLM=GPT-5
2026.06
70.69
CoT
Judge Model=GPT-4o
2026.06
70.69
ToT
Judge Model=GPT-4o
2026.06
69.28
DICAI
Judge Model=GPT-4o
2026.06
68.7
CoT-SC
Judge LLM=GPT-5
2026.06
68.43
Self-Refine
Judge LLM=GPT-5
2026.06
68.02
ICAI
Judge LLM=GPT-5
2026.06
66.8
CoT-SC
Judge Model=GPT-4o
2026.06
66.15
AutoRubric
Judge LLM=GPT-5
2026.06
63.21
ICAI
Judge Model=GPT-4o
2026.06
62.89
AutoRubric
Judge Model=GPT-4o
2026.06
59.22
Feedback
Search any
task
Search any
task