arrow
Return

Constrained continuous-action reinforcement learning for supply chain inventory management

delete2024-02-01
delete7
delete
OA
AI
R
Radu Burtea
C
Calvin Tsay *
DOI:10.1016/j.compchemeng.2023.108518delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Reinforcement learning (RL) is a promising solution for difficult decision-making problems, such as inventory management in chemical supply chains. However, enabling RL to explicitly consider known environment constraints is crucial for safe deployment in practical applications. This work incorporates recent tools for optimization over trained neural networks to introduce two algorithms for safe training and deployment of RL, with a focus on supply chains. Specifically, we use optimization over trained neural-network state-action value functions (i.e., a critic function) to directly incorporate constraints when computing actions in a continuous action space. Furthermore, we introduce a second algorithm that guarantees constraint satisfaction during deployment by directly implementing actions from constrained optimization of a trained value function. The algorithms are compared against state-of-the-art algorithms TRPO, CPO, and RCPO using a computational supply chain case study.
Keywords:
Safe reinforcement learning
Optimization and machine learning toolkit
(OMLT)
Continuous action Q-learning
Inventory management problem
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

C
Computers and Chemical Engineering
IF:
3.9
Papers:
8.1K
Citations:
1.7W

Organization

I
Imperial College London
Scholars:
8.3W
Papers: 7.3W
Citations: 11.1W