arrow
Return

Entropy regularized reinforcement learning using large deviation theory

delete2023-05-10
delete2
delete
OA
AI
A
Argenis Arriojas *
J
Jacob Adamczyk
S
Stas Tiomkin
R
Rahul V. Kulkarni
DOI:10.1103/PhysRevResearch.5.023085delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Reinforcement learning (RL) is an important field of research in machine learning that is increasingly being applied to complex optimization problems in physics. In parallel, concepts from physics have contributed to important advances in RL with developments such as entropy-regularized RL. While these developments have led to advances in both fields, obtaining analytical solutions for optimization in entropy-regularized RL is currently an open problem. In this paper, we establish a mapping between entropy-regularized RL and research in nonequilibrium statistical mechanics focusing on Markovian processes conditioned on rare events. In the long-time limit, we apply approaches from large deviation theory to derive exact analytical results for the optimal policy and optimal dynamics in Markov decision process (MDP) models of reinforcement learning. The results obtained lead to an analytical and computational framework for entropy-regularized RL which is validated by simulations. The mapping established in this work connects current research in reinforcement learning and nonequilibrium statistical mechanics, thereby opening avenues for the application of analytical and computational approaches from one field to cutting-edge problems in the other.
Keywords:
DEEP

Journal

Physical Review Research cover
Physical Review Research
IF:
4.2
Papers:
7.6K
Citations:
2.7W

Organization

U
university of massachusetts system
Scholars:
3.8W
Papers: 3.5W
Citations: 42
U
University of Massachusetts Boston
Scholars:
2.4K
Papers: 1.9K
Citations: 4.2K
Cited Papers

Cited Papers

errShare
errSave
Local vibrational modes in Mg-doped gallium nitride
err1994-05-15
err0
PREAI
errM. S. Brandt; J. W. Ager; W. Götz; N. M. Johnson; J. S. Harris; R. J. Molnar; T. D. Moustakas
errShare
errSave
Evolutionary reinforcement learning of dynamical large deviations
err2020-07-27
err23
errOAAI
errWhitelam, Stephen; Jacobson, Daniel; Tamblyn, Isaac
errShare
errSave
errShare
errSave
errShare
errSave
Customization of Avatars in a HPV Digital Gaming Intervention for College-Age Males: An Experimental Study
err2018-10-17
err0
PREAI
errGabrielle Darville; Charkarra Anderson – Lewis; Michael Stellefson; Yu-Hao Lee; Jann MacInnes; R. Morgan Pigg; Juan E. Gilbert; Sanethia Thomas
errShare
errSave
researcher View more