Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multi-turn Refinement on DeepResearch Bench 100 follow-up instances
Loading...
15.35
Instance Score
SCAFFOLDAGENT
12.542
13.271
14
14.729
Jun 18, 2026
Instance Score
Evidence Score
Logic Score
Completeness Score
Explanation Score
Total Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
Instance Score
Evidence Score
Logic Score
Completeness Score
Explanation Score
Total Score
SCAFFOLDAGENT
Setting=SCAFFOLDAGENT
2026.06
15.35
13.4
14.55
15.1
14.2
72.6
ReAct-FW
Setting=ReAct-FW
2026.06
13.8
12.2
12.75
6.6
11.41
57.76
ReAct-GR
Setting=ReAct-GR
2026.06
12.65
10.85
11.55
10.9
13.75
59.7
Feedback
Search any
task
Search any
task