Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Scientific Reasoning on ScienceWorld Unseen
Loading...
62
Average Reward
Co-Evolving Agents
-1.762608
14.791146
31.3449
47.898654
Nov 27, 2025
Dec 28, 2025
Jan 28, 2026
Feb 28, 2026
Mar 31, 2026
May 1, 2026
Jun 2, 2026
Average Reward
Updated 1mo ago
Evaluation Results
Method
Method
Links
Average Reward
Co-Evolving Agents
Adaptation=Fine-tuning...
2025.11
62
Co-Evolving Agents
Backbone=Qwen3-4B-Inst...
2025.11
58.5
Llama-2-7B-Chat + ETO
Adaptation=Fine-tuning...
2025.11
55.5
ETO
Backbone=Qwen3-4B-Inst...
2025.11
55.2
Llama-2-7B-Chat + RFT
Adaptation=Fine-tuning...
2025.11
54.3
Llama-2-7B-Chat + PPO
Adaptation=Fine-tuning...
2025.11
51.7
Llama-2-7B-Chat + SFT
Adaptation=Fine-tuning...
2025.11
41.9
SFT
Backbone=Qwen3-4B-Inst...
2025.11
40.8
GPT-4
Adaptation=In-context
2025.11
38.1
GPT-3.5-Turbo
Adaptation=In-context
2025.11
10.5
DELTAMEM
LLM=DeepSeek-V4-flash,...
2026.06
0.8688
Synapse
LLM=DeepSeek-V4-flash,...
2026.06
0.8558
AWM
LLM=DeepSeek-V4-flash,...
2026.06
0.7453
No Memory
LLM=DeepSeek-V4-flash,...
2026.06
0.7186
RBank
LLM=DeepSeek-V4-flash,...
2026.06
0.6898
Feedback
Search any
task
Search any
task