Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Mixed-Motive Multi-Agent Reinforcement Learning on MPE simple_tag
Loading...
227.69
Reward
MAPPO
-2.0044
57.6278
117.26
176.8922
Jun 9, 2026
Reward
Cumulative Regret
Regret Gap
Updated 1mo ago
Evaluation Results
Method
Method
Links
Reward
Cumulative Regret
Regret Gap
MAPPO
2026.06
227.69
1,084.68
2,192.3
Φ-AC
2026.06
12.55
481.56
53.79
MADDPG
2026.06
6.83
1,409.98
2,321.88
Feedback
Search any
task
Search any
task