arrow
Return

Deep Policy Iteration with Integer Programming for Inventory Management

delete2025-01-06
delete0
delete
OA
AI
P
Pavithra Harsha *
J
Jayant Kalagnanam
B
Brian Quanz
D
Divya Singhvi *
DOI:10.1287/msom.2022.0617delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Problem definition: In this paper, we present a reinforcement learning (RL)based framework for optimizing long-term discounted reward problems with large combinatorial action space and state dependent constraints. These characteristics are common to many operations management problems, for example, network inventory replenishment, where managers have to deal with uncertain demand, lost sales, and capacity constraints that results in more complex feasible action spaces. Our proposed programmable actor RL (PARL) uses a deep-policy iteration method that leverages neural networks to approximate the value function and combines it with mathematical programming and sample average approximation to solve the per-step-action optimally while accounting for combinatorial action spaces and state-dependent constraint sets. Methodology/results: We then show how the proposed methodology can be applied to complex inventory replenishment problems where analytical solutions are intractable. We also benchmark the proposed algorithm against state-of-the-art RL algorithms and commonly used replenishment heuristics and find that the proposed algorithm considerably outperforms existing methods by as much as 14.7% on average in various complex supply chain settings. Managerial implications: We find that this improvement in performance of PARL over benchmark algorithms can be directly attributed to better inventory cost management, especially in inventory constrained settings. Furthermore, in the simpler setting where optimal replenishment policy is tractable or known near optimal heuristics exist, we find that the RL-based policies can learn near optimal policies. Finally, to make RL algorithms more accessible for inventory management researchers, we also discuss the development of a modular Python library that can be used to test the performance of RL algorithms with various supply chain structures. This library can spur future research in developing practical and near-optimal algorithms for inventory management problems.
Keywords:
multiechelon inventory management
inventory replenishment
deep reinforcement learning

Journal

Manufacturing and Service Operations Management cover
Manufacturing and Service Operations Management
IF:
4.2
Papers:
393
Citations:
7.1K

Organization

N
New York University
Scholars:
4.4W
Papers: 3.9W
Citations: 5.8W
I
international business machines (ibm)
Scholars:
5.7K
Papers: 4.5K
Citations: 4
I
ibm usa
Scholars:
1.4K
Papers: 1.0K
Citations: 0
researcher View more organizations