返回
Automatic Temperature Parameter Tuning for Reinforcement Learning Using Path Integral Policy Improvement
DOI:10.1109/TNNLS.2023.3312857.png)
摘要
En 中文
In this article, we propose a novel variant of path integral policy improvement with covariance matrix adaptation(PI2-CMA), which is a reinforcement learning (RL) algorithm that aims to optimize a parameterized policy for the continuous behavior of robots. PI2-CMA has a hyperparameter called the temperature parameter, and its value is critical for performance; however, little research has been conducted on it and the existing method still contains a tunable parameter, which can be critical to performance. Therefore, tuning by trial and error is necessary in the existing method. Moreover, we show that there is a problem setting that cannot be learned by the existing method. The pro-posed method solves both problems by automatically adjusting the temperature parameter for each update. We confirmed the effectiveness of the proposed method using numerical tests.
Keyword:
Legged robot
policy improvement
reinforcement learning (RL)
robotics
snake robot
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
A Path-Integral-Based Reinforcement Learning Algorithm for Path Following of an Autoassembly Mobile Robot一种基于路径积分的强化学习算法,用于自动装配移动机器人的路径跟踪

