Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PROPEL: Supervised and Reinforcement Learning for Large-Scale Supply Chain Planning

About

This paper considers how to fuse Machine Learning (ML) and optimization to solve large-scale Supply Chain Planning (SCP) optimization problems. These problems can be formulated as MIP models which feature both integer (non-binary) and continuous variables, as well as flow balance and capacity constraints. This raises fundamental challenges for existing integrations of ML and optimization that have focused on binary MIPs and graph problems. To address these, the paper proposes PROPEL, a new framework that combines optimization with both supervised and Deep Reinforcement Learning (DRL) to reduce the size of search space significantly. PROPEL uses supervised learning, not to predict the values of all integer variables, but to identify the variables that are fixed to zero in the optimal solution, leveraging the structure of SCP applications. PROPEL includes a DRL component that selects which fixed-at-zero variables must be relaxed to improve solution quality when the supervised learning step does not produce a solution with the desired optimality tolerance. PROPEL has been applied to industrial supply chain planning optimizations with millions of variables. The computational results show dramatic improvements in solution times and quality, including a 60% reduction in primal integral and an 88% primal gap reduction, and improvement factors of up to 13.57 and 15.92, respectively.

Vahid Eghbal Akhlaghi, Reza Zandehshahvar, Pascal Van Hentenryck• 2025

Related benchmarks

TaskDatasetResultRank
Large-scale OptimizationOCP Hard
PG Mean (%)8
4
Large-scale OptimizationOCP Very-Hard
PG Mean1.72
4
Large-scale OptimizationMMCNP Hard
PG Mean9
4
Large-scale OptimizationMMCNP Very-Hard
PG Mean (%)0.15
4
Large-scale OptimizationSLAP Hard
PG Mean2.8
4
Large-scale OptimizationSLAP Very-Hard
Performance Gap (PG) Mean15
4
Large-scale OptimizationCOURSE Hard
PG Mean1.3
4
Large-scale OptimizationCOURSE Very-Hard
Performance Gap (Mean)13
4
Large-scale OptimizationOFP Hard
PG Mean (%)20
3
Large-scale OptimizationOFP Very-Hard
Performance Gap (PG) Mean0.78
3
Showing 10 of 10 rows

Other info

Follow for update