Preference-Driven Multi-Objective Combinatorial Optimization with Conditional Computation

About

Recent deep reinforcement learning methods have achieved remarkable success in solving multi-objective combinatorial optimization problems (MOCOPs) by decomposing them into multiple subproblems, each associated with a specific weight vector. However, these methods typically treat all subproblems equally and solve them using a single model, hindering the effective exploration of the solution space and thus leading to suboptimal performance. To overcome the limitation, we propose POCCO, a novel plug-and-play framework that enables adaptive selection of model structures for subproblems, which are subsequently optimized based on preference signals rather than explicit reward values. Specifically, we design a conditional computation block that routes subproblems to specialized neural architectures. Moreover, we propose a preference-driven optimization algorithm that learns pairwise preferences between winning and losing solutions. We evaluate the efficacy and versatility of POCCO by applying it to two state-of-the-art neural methods for MOCOPs. Experimental results across four classic MOCOP benchmarks demonstrate its significant superiority and strong generalization.

Mingfeng Fan, Jianan Zhou, Yifeng Zhang, Yaoxin Wu, Jinbiao Chen, Guillaume Adrien Sartoretti• 2025

Related benchmarks

Task	Dataset	Result
Bi-objective Traveling Salesman Problem	Bi-TSP50	Hypervolume (HV)0.6418	44
Multi-Objective Traveling Salesperson Problem	KroAB200	Hypervolume (HV)73.69	44
Multi-Objective Traveling Salesperson Problem	KroAB100	Hypervolume (HV)0.7006	44
Tri-Objective Traveling Salesman Problem	Tri-TSP50	Hypervolume (HV)0.4437	44
Multi-Objective Traveling Salesperson Problem	KroAB150	Hypervolume (HV)69.76	44
Multi-objective Knapsack Problem	Bi-KP n=50	HV0.3562	34
Multi-objective Knapsack Problem	Bi-KP n=100	HV0.4535	34
Multi-objective Knapsack Problem	Bi-KP n=200	HV0.3603	34
Bi-objective Traveling Salesman Problem	Bi-TSP20	Hypervolume (HV)0.6275	24
Bi-objective Traveling Salesman Problem	Bi-TSP 100	Hypervolume (HV)70.77	24

Showing 10 of 30 rows

Other info

Follow for update

@wizwand_team Discord