arrow
返回

V-learning: An adaptive reinforcement learning algorithm for the optimal stopping problem

delete2023-11-01
delete1
PRE
AI
C
Chi-Guhn Lee *
DOI:10.1016/j.eswa.2023.120702delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The optimal stopping problem is concerned with finding an optimal policy to stop a stochastic process in order to maximize the expected return. This problem is critical in stochastic control and can be found in many different fields, such as operations research, finance, and healthcare. In this paper, we model the underlying stochastic process of the optimal stopping problem as a Markov decision process and propose a computationally efficient model-free value-based reinforcement learning approach, named ������V-learning. The efficiency is improved by taking advantage of the unique structural properties of the optimal stopping problem into our algorithm design. We consider two types of the optimal stopping problems: the standard optimal stopping and the regenerative optimal stopping, which differ in their transition dynamics once the stopping action is executed. We conduct numerical experiments on our proposed method and compare its performance against existing reinforcement learning algorithms and rule-based policies. The results show that our ������V-learning method is able to outperform the benchmark algorithms in all experiments.
Keyword:
Reinforcement learning
Optimal stopping
Markov decision process
Stochastic control
Secretary problem
Optimal replacement

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
3.0W
被引数:
10.2W

机构

U
university of toronto
学者数:
14.8W
论文数: 12.0W
被引数: 165
引用论文

引用论文

Poly(3,4‐ethylenedioxythiophene): Chemical Synthesis, Transport Properties, and Thermoelectric Devices
err2019-03-18
err0
PREAI
errIoannis Petsagkourakis; Nara Kim; Klas Tybrandt; Igor Zozoulenko; Xavier Crispin
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
Model-based Reinforcement Learning: A Survey基于模型的强化学习: 综述
err2023-01-01
err161
errOAAI
errMoerland, Thomas M.; Broekens, Joost; Plaat, Aske; Jonker, Catholijn M.
err分享
err收藏
学者 查看更多内容