Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
About
For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than relying on complex reward functions and explicit cooperation mechanisms, we ask what minimal ingredients are required for effective coordination and exploration to emerge in multi-agent settings. We investigate this question through self-supervised goal-reaching, where agents aim to maximize the likelihood of visiting a goal state rather than maximizing a reward. Despite a sparse feedback signal, we present empirical results that show self-supervised goal-reaching techniques enable agents to learn from such feedback. On MARL benchmarks, self-supervised goal-reaching outperforms alternative approaches that have access to the same sparse reward signal. Furthermore, we empirically demonstrate that multi-agent self-supervised goal-reaching approaches can be more robust than single-agent strategies. While there is no explicit exploration mechanism, this approach explores nontrivial intermediate coordination strategies in sparse settings where alternative approaches fail to achieve a single success.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-agent cooperation | Simple_Tag 6 agents | Average Reward1.63e+4 | 44 | |
| 2s3z | StarCraft Multi-Agent Challenge Smax | Mean Win Rate95 | 5 | |
| 3m | StarCraft Multi-Agent Challenge Smax | Mean Win Rate94 | 3 | |
| 3s_v_5z | StarCraft Multi-Agent Challenge Smax | Mean Win Rate95 | 3 | |
| 6h_v_8z | StarCraft Multi-Agent Challenge Smax | Mean Win Rate100 | 3 | |
| 8m | StarCraft Multi-Agent Challenge Smax | Mean Win Rate84 | 3 | |
| Ant | Multi-Agent Control | Mean Success Rate270.6 | 3 | |
| Humanoid | Multi-Agent Control | Mean Success Rate3.38 | 3 | |
| Half-Cheetah | Multi-Agent Control | Mean Success Rate901.6 | 3 | |
| 3 Agents | MPE Tag | Mean Episode Return4.70e+3 | 2 |