返回
V-learning: An adaptive reinforcement learning algorithm for the optimal stopping problem
DOI:10.1016/j.eswa.2023.120702.png)
摘要
En 中文
The optimal stopping problem is concerned with finding an optimal policy to stop a stochastic process in order to maximize the expected return. This problem is critical in stochastic control and can be found in many different fields, such as operations research, finance, and healthcare. In this paper, we model the underlying stochastic process of the optimal stopping problem as a Markov decision process and propose a computationally efficient model-free value-based reinforcement learning approach, named ������V-learning. The efficiency is improved by taking advantage of the unique structural properties of the optimal stopping problem into our algorithm design. We consider two types of the optimal stopping problems: the standard optimal stopping and the regenerative optimal stopping, which differ in their transition dynamics once the stopping action is executed. We conduct numerical experiments on our proposed method and compare its performance against existing reinforcement learning algorithms and rule-based policies. The results show that our ������V-learning method is able to outperform the benchmark algorithms in all experiments.
Keyword:
Reinforcement learning
Optimal stopping
Markov decision process
Stochastic control
Secretary problem
Optimal replacement
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
引用论文
Synthesis, structure, and characterization of a new promising nonlinear optical crystal: Cd5(BO3)3F
CrystEngComm
IF0
Fabrication of porous hollow γ-Al2O3 nanofibers by facile electrospinning and its application for water remediation静电纺丝法制备多孔中空 γ-Al2O3纳米纤维及其在水体修复中的应用

