arrow
返回

Deep Policy Iteration with Integer Programming for Inventory Management

delete2025-01-06
delete0
delete
OA
AI
P
Pavithra Harsha *
J
Jayant Kalagnanam
B
Brian Quanz
D
Divya Singhvi *
DOI:10.1287/msom.2022.0617delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Problem definition: In this paper, we present a reinforcement learning (RL)based framework for optimizing long-term discounted reward problems with large combinatorial action space and state dependent constraints. These characteristics are common to many operations management problems, for example, network inventory replenishment, where managers have to deal with uncertain demand, lost sales, and capacity constraints that results in more complex feasible action spaces. Our proposed programmable actor RL (PARL) uses a deep-policy iteration method that leverages neural networks to approximate the value function and combines it with mathematical programming and sample average approximation to solve the per-step-action optimally while accounting for combinatorial action spaces and state-dependent constraint sets. Methodology/results: We then show how the proposed methodology can be applied to complex inventory replenishment problems where analytical solutions are intractable. We also benchmark the proposed algorithm against state-of-the-art RL algorithms and commonly used replenishment heuristics and find that the proposed algorithm considerably outperforms existing methods by as much as 14.7% on average in various complex supply chain settings. Managerial implications: We find that this improvement in performance of PARL over benchmark algorithms can be directly attributed to better inventory cost management, especially in inventory constrained settings. Furthermore, in the simpler setting where optimal replenishment policy is tractable or known near optimal heuristics exist, we find that the RL-based policies can learn near optimal policies. Finally, to make RL algorithms more accessible for inventory management researchers, we also discuss the development of a modular Python library that can be used to test the performance of RL algorithms with various supply chain structures. This library can spur future research in developing practical and near-optimal algorithms for inventory management problems.
Keyword:
multiechelon inventory management
inventory replenishment
deep reinforcement learning

期刊

Manufacturing and Service Operations Management 封面图
Manufacturing and Service Operations Management
IF:
4.2
论文数:
393
被引数:
7.1K

机构

N
New York University
学者数:
4.4W
论文数: 3.9W
被引数: 5.8W
I
international business machines (ibm)
学者数:
5.7K
论文数: 4.5K
被引数: 4
I
ibm usa
学者数:
1.4K
论文数: 1.0K
被引数: 0
学者 查看更多机构
引用论文

引用论文

TOMATO GROWTH AND YIELD AFFECTED BY NICKEL PRESENTED IN THE NUTRIENT SOLUTION
err1998-04-01
err0
PREAI
errJ. Balaguer; M.B. Almendro; I. Gómez; J. Navarro Pedreño; J. Mataix
err分享
err收藏
Optimales Thermomanagement und Elektrifizierung in 48-V-Hybriden
err2018-09-10
err0
PREAI
errFriedrich Graf; Stefan Lauer; Johannes Hofstetter; Mattia Perugini
err分享
err收藏
The L‐arginine inhibition of rat middle cerebral artery contractile responses is mediated by inducible nitric oxide synthase
err2008-10-09
err0
PREAI
errM. J. Alonso; M. A. Rodríguez‐Martínez; J. Martínez‐Orgado; J. Marín; M. Salaices
err分享
err收藏
A typology and literature review on stochastic multi-echelon inventory models随机多级库存模型的类型学及文献综述
err2018-09-01
err81
errOAAI
errde Kok, Ton; Grob, Christopher; Laumanns, Marco; Minner, Stefan; Rambau, Joerg; Schade, Konrad
err分享
err收藏
Deep reinforcement learning for inventory control: A roadmap
err2022-04-01
err78
errOAAI
errBoute, Robert N.; Gijsbrechts, Joren; van Jaarsveld, Willem; Vanvuchelen, Nathalie
err分享
err收藏
学者 查看更多内容