arrow
Return

A Generalized Stacked Reinforcement Learning Method for Sampled Systems

delete2023-11-01
delete6
PRE
AI
P
Pavel Osinenko *
D
Dmitrii Dobriborsci
G
Grigory Yaremenko
G
Georgiy Malaniya
DOI:10.1109/TAC.2023.3250032delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
A common setting of reinforcement learning (RL) is a Markov decision process (MDP) in which the environment is a stochastic discrete-time dynamical system. Whereas MDPs are suitable in such applications as video games or puzzles, physical systems are time continuous. A general variant of RL is of digital format, where updates of the value (or cost) and policy are performed at discrete moments in time. The agent-environment loop then amounts to a sampled system, whereby sample-and-hold is a specific case. In this article, we propose and benchmark two RL methods suitable for sampled systems. Specifically, we hybridize model predictive control with critics learning the optimal Q- and value (or cost-to-go) function. Optimality is analyzed and performance comparison is done in an experimental case study with a mobile robot.
Keywords:
Mobile robot
model predictive control (MPC)
optimal control
q-learning
reinforcement learning (RL)

Journal

IEEE Transactions on Automatic Control cover
IEEE Transactions on Automatic Control
IF:
7
Papers:
1.3W
Citations:
6.7W

Organization

S
skolkovo institute of science & technology
Scholars:
3.3K
Papers: 2.3K
Citations: 1