Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning

About

Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives. In this setting, conflicts arise not only across objectives but also across agents with different observations, roles, and contributions. We propose Preference Coordinated Multi-agent Policy Optimization (PCMA), which learns coordinated agent-specific preferences to enable complementary trade-offs among agents. Theoretically, we formulate cooperative MOMARL as a team-optimal game and show that, under suitable conditions, preference diversity can induce team improvement through a first-order improvement decomposition. Experiments on multiple cooperative MOMA environments and a practical traffic-control scenario show that PCMA improves both performance and trade-off coordination.

Pengxin Wang, Lihao Guo, Yi Xie, Bo Liu, Siyang Cao, Jingdi Chen• 2026

Related benchmarks

TaskDatasetResultRank
Multi-Agent Reinforcement LearningSMAC 2s3z
Win Rate100
7
Cooperative Multi-Agent TaskCooperative Spread
Success Rate100
4
Cooperative Multi-Agent TaskSafe Predator Prey
Success Rate96
4
Cooperative Multi-Agent TaskCATCH
Success Rate94
4
Cooperative Multi-Agent TaskEscort
Average Reward17.29
4
Cooperative Multi-Agent TaskMOMAwalker
Forward Distance93.64
4
Cooperative Multi-Agent TaskSMAC 8m
Success Rate87
4
Cooperative Multi-Agent TaskSMAC 3m
Success Rate97
4
Preference-conditioned controlOpenCDA-MARL CARLA Cooperative
Utility-2.07e+3
3
Preference-conditioned controlOpenCDA-MARL CARLA Competitive
Utility-2.88e+3
3
Showing 10 of 10 rows

Other info

Follow for update