arrow
返回

Robust Q-learning algorithm for Markov decision processes under Wasserstein uncertainty

delete2024-10-01
delete2
PRE
AI
A
Ariel Neufeld *
J
Julian Sester
DOI:10.1016/j.automatica.2024.111825delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We present a novel Q-learning algorithm tailored to solve distributionally robust Markov decision problems where the corresponding ambiguity set of transition probabilities for the underlying Markov decision process is a Wasserstein ball around a (possibly estimated) reference measure. We prove convergence of the presented algorithm and provide several examples also using real data to illustrate both the tractability of our algorithm as well as the benefits of considering distributional robustness when solving stochastic optimal control problems, in particular when the estimated distributions turn out to be misspecified in practice. (c) 2024 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
Keyword:
Markov decision process
Wasserstein uncertainty
Distributionally robust optimization
Reinforcement learning
Q-learning

期刊

Automatica 封面图
Automatica
IF:
5.9
论文数:
1.2W
被引数:
5.2W

机构

N
Nanyang Technological University
学者数:
4.9W
论文数: 4.8W
被引数: 8.1W
N
National University of Singapore
学者数:
7.6W
论文数: 6.5W
被引数: 11.4W
引用论文

引用论文

Topics in Polymer Physics
err
IF0
err2006-10-24
err0
PREAI
errRichard S Stein; Joseph Powers
err分享
err收藏
Single Event Effects Characterization of the Programmable Logic of Xilinx Zynq-7000 FPGA Using Very/Ultra High-Energy Heavy Ions
err2021-01-01
err0
PREAI
errVasileios Vlagkoulis; Aitzan Sari; John Vrachnis; Georgios Antonopoulos; Nikolaos Segkos; Mihalis Psarakis; Antonios Tavoularis; Gianluca Furano; Cesar Boatella Polo; Christian Poivey; Veronique Ferlet-Cavrois; Maria Kastriotou; Pablo Fernandez Martinez; Ruben Garcia Alia; Kay-Obbe Voss; Christoph Schuy
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容