Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability
About
Many real-world tasks involve multiple agents with partial observability and limited communication. Learning is challenging in these settings due to local viewpoints of agents, which perceive the world as non-stationary due to concurrently-exploring teammates. Approaches that learn specialized policies for individual tasks face problems when applied to the real world: not only do agents have to learn and store distinct policies for each task, but in practice identities of tasks are often non-observable, making these approaches inapplicable. This paper formalizes and addresses the problem of multi-task multi-agent reinforcement learning under partial observability. We introduce a decentralized single-task learning approach that is robust to concurrent interactions of teammates, and present an approach for distilling single-task policies into a unified policy that performs well across multiple related tasks, without explicit provision of task identity.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-Agent Reinforcement Learning | SMAC 3svs3z to 3svs4z | Environment Steps to 80% Reward2.59 | 7 | |
| Multi-Agent Reinforcement Learning | SMAC 3svs3z to 3svs5z | Environment Steps to 80% Reward3.42 | 7 | |
| Multi-Agent Reinforcement Learning | SMAC 8mvs8m to 8mvs9m | Steps to 80% Reward4.28 | 7 | |
| Multi-Agent Transfer Learning | SMAC 3svs3z to 3svs5z | Final Win Rate92 | 7 | |
| Multi-Agent Transfer Learning | SMAC 3svs3z to 3svs4z | Final Win Rate0.86 | 7 | |
| Multi-Agent Transfer Learning | SMAC 8mvs8m to 8mvs9m | Final Win Rate86 | 7 |