Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Science Reasoning and Simulation on ScienceWorld In-domain
Loading...
62
Score
Task-Perturbed NLL Optimization
45.36
49.68
54
58.32
Jun 25, 2026
Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
Score
Task-Perturbed NLL Optimization
Backbone=Qwen3-8B
2026.06
62
GRPO
Backbone=Qwen3-8B
2026.06
57
Task-Perturbed NLL Optimization
Backbone=Qwen3-4B
2026.06
51
GRPO
Backbone=Qwen3-4B
2026.06
46
Feedback
Search any
task
Search any
task