Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning
About
Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives. In this setting, conflicts arise not only across objectives but also across agents with different observations, roles, and contributions. We propose Preference Coordinated Multi-agent Policy Optimization (PCMA), which learns coordinated agent-specific preferences to enable complementary trade-offs among agents. Theoretically, we formulate cooperative MOMARL as a team-optimal game and show that, under suitable conditions, preference diversity can induce team improvement through a first-order improvement decomposition. Experiments on multiple cooperative MOMA environments and a practical traffic-control scenario show that PCMA improves both performance and trade-off coordination.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-Agent Reinforcement Learning | SMAC 2s3z | Win Rate100 | 7 | |
| Cooperative Multi-Agent Task | Cooperative Spread | Success Rate100 | 4 | |
| Cooperative Multi-Agent Task | Safe Predator Prey | Success Rate96 | 4 | |
| Cooperative Multi-Agent Task | CATCH | Success Rate94 | 4 | |
| Cooperative Multi-Agent Task | Escort | Average Reward17.29 | 4 | |
| Cooperative Multi-Agent Task | MOMAwalker | Forward Distance93.64 | 4 | |
| Cooperative Multi-Agent Task | SMAC 8m | Success Rate87 | 4 | |
| Cooperative Multi-Agent Task | SMAC 3m | Success Rate97 | 4 | |
| Preference-conditioned control | OpenCDA-MARL CARLA Cooperative | Utility-2.07e+3 | 3 | |
| Preference-conditioned control | OpenCDA-MARL CARLA Competitive | Utility-2.88e+3 | 3 |