Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
In-context adaptation on Mastermind
Loading...
7.13
Cumulative Reward
Cross-task RL
4.2076
4.9663
5.725
6.4837
Jun 13, 2026
Cumulative Reward
Gain
Final Reward
Updated 1mo ago
Evaluation Results
Method
Method
Links
Cumulative Reward
Gain
Final Reward
Cross-task RL
Training Strategy=Cros...
2026.06
7.13
3
0.7
Single-task RL
Training Strategy=Sing...
2026.06
5.12
18
0.51
Base (Qwen3-8B)
Training Strategy=None...
2026.06
4.32
4
0.44
Feedback
Search any
task
Search any
task