返回
Actor-Critic Learning Control Based on l2-Regularized Temporal-Difference Prediction With Gradient Correction
DOI:10.1109/TNNLS.2018.2808203.png)
摘要
En 中文
Actor-critic based on the policy gradient (PG-based AC) methods have been widely studied to solve learning control problems. In order to increase the data efficiency of learning prediction in the critic of PG-based AC, studies on how to use recursive least-squares temporal difference (RLS-TD) algorithms for policy evaluation have been conducted in recent years. In such contexts, the critic RLS-TD evaluates an unknown mixed policy generated by a series of different actors, but not one fixed policy generated by the current actor. Therefore, this AC framework with RLS-TD critic cannot be proved to converge to the optimal fixed point of learning problem. To address the above problem, this paper proposes a new AC framework named critic-iteration PG (CIPG), which learns the state-value function of current policy in an on-policy way and performs gradient ascent in the direction of improving discounted total reward. During each iteration, CIPG keeps the policy parameters fixed and evaluates the resulting fixed policy by l(2)-regularized RLS-TD critic. Our convergence analysis extends previous convergence analysis of PG with function approximation to the case of RLS-TD critic. The simulation results demonstrate that the l(2)-regularization term in the critic of CIPG is undamped during the learning process, and CIPG has better learning efficiency and faster convergence rate than conventional AC learning control methods.
Keyword:
l(2)-regularization
actor-critic (AC)
policy gradient (PG)
reinforcement learning (RL)
value function approximation
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
A convergent actor-critic-based FRL algorithm with application to power management of wireless transmitters基于actor-critic的收敛FRL算法及其在无线发射机电源管理中的应用

