arrow
返回

Parameterized MDPs and Reinforcement Learning Problems--A Maximum Entropy Principle-Based Framework

delete2022-09-01
delete4
delete
OA
AI
A
Amber Srivastava *
S
Srinivasa M. Salapaka
DOI:10.1109/TCYB.2021.3102510delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
We present a framework to address a class of sequential decision-making problems. Our framework features learning the optimal control policy with robustness to noisy data, determining the unknown state and action parameters, and performing sensitivity analysis with respect to problem parameters. We consider two broad categories of sequential decision-making problems modeled as infinite horizon Markov decision processes (MDPs) with (and without) an absorbing state. The central idea underlying our framework is to quantify exploration in terms of the Shannon entropy of the trajectories under the MDP and determine the stochastic policy that maximizes it while guaranteeing a low value of the expected cost along a trajectory. This resulting policy enhances the quality of exploration early on in the learning process, and consequently allows faster convergence rates and robust solutions even in the presence of noisy data as demonstrated in our comparisons to popular algorithms, such as Q-learning, Double Q-learning, and entropy regularized Soft Q-learning. The framework extends to the class of parameterized MDP and RL problems, where states and actions are parameter dependent, and the objective is to determine the optimal parameters along with the corresponding optimal policy. Here, the associated cost function can possibly be nonconvex with multiple poor local minima. Simulation results applied to a 5G small cell network problem demonstrate the successful determination of communication routes and the small cell locations. We also obtain sensitivity measures to problem parameters and robustness to noisy environment data.
Keyword:
Entropy
Cost function
Noise measurement
5G mobile communication
Heuristic algorithms
Decision making
Convergence
Markov decision processes (MDPs)
maximum entropy principle (MEP)
network design
parameterized sequential decision making
reinforcement learning

期刊

IEEE Transactions on Cybernetics 封面图
IEEE Transactions on Cybernetics
IF:
10.5
论文数:
1.1W
被引数:
5.0W

机构

University of Illinois System 封面图
University of Illinois System
学者数:
6.8W
论文数: 6.2W
被引数: 644
引用论文

引用论文

Platelets and Related Products
err2007-01-01
err0
PREAI
errJohn M. Fisk; Patricia T. Pisciotto; Edward L. Snyder; Peter L. Perrotta
err分享
err收藏
err分享
err收藏
Ion-specificity and surface water dynamics in protein solutions
err2018-01-01
err0
errOAAI
errTadeja Janc; Miha Lukšič; Vojko Vlachy; Baptiste Rigaud; Anne-Laure Rollet; Jean-Pierre Korb; Guillaume Mériguet; Natalie Malikova
err分享
err收藏
A comparative analysis of binding in ultralong-range Rydberg molecules
err2015-05-07
err0
errOAAI
errC Fey; M Kurz; P Schmelcher; S T Rittenhouse; H R Sadeghpour
err分享
err收藏
err分享
err收藏
What Is the Relationship of Fear Avoidance to Physical Function and Pain Intensity in Injured Athletes?
err2018-02-16
err0
errOAAI
errStefan F. Fischerauer; Mojtaba Talaei-Khoei; Rens Bexkens; David C. Ring; Luke S. Oh; Ana-Maria Vranceanu
err分享
err收藏
err分享
err收藏
Locally Weighted Ensemble Clustering
err2018-05-01
err282
errOAAI
errHuang, Dong; Wang, Chang-Dong; Lai, Jian-Huang
err分享
err收藏
学者 查看更多内容