arrow
返回

Actor-Critic Learning Control With Regularization and Feature Selection in Policy Gradient Estimation

delete2021-03-01
delete18
PRE
AI
L
Luntong Li
D
Dazi Li *
T
Tianheng Song
徐鑫 封面图
徐鑫 (Xin Xu)
DOI:10.1109/TNNLS.2020.2981377delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Actor-critic (AC) learning control architecture has been regarded as an important framework for reinforcement learning (RL) with continuous states and actions. In order to improve learning efficiency and convergence property, previous works have been mainly devoted to solve regularization and feature learning problem in the policy evaluation. In this article, we propose a novel AC learning control method with regularization and feature selection for policy gradient estimation in the actor network. The main contribution is that l(1)-regularization is used on the actor network to achieve the function of feature selection. In each iteration, policy parameters are updated by the regularized dual-averaging (RDA) technique, which solves a minimization problem that involves two terms: one is the running average of the past policy gradients and the other is the l(1)-regularization term of policy parameters. Our algorithm can efficiently calculate the solution of the minimization problem, and we call the new adaptation of policy gradient RDApolicy gradient (RDA-PG). The proposed RDA-PG can learn stochastic and deterministic near-optimal policies. The convergence of the proposed algorithm is established based on the theory of two-timescale stochastic approximation. The simulation and experimental results show that RDA-PG performs feature selection successfully in the actor and learns sparse representations of the actor both in stochastic and deterministic cases. RDA-PG performs better than existing AC algorithms on standard RL benchmark problems with irrelevant features or redundant features.
Keyword:
l(1)-regularization
actor-critic (AC)
policy gradient
regularized dual-averaging (RDA)
reinforcement learning (RL)
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

IEEE Transactions on Neural Networks and Learning Systems 封面图
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
论文数:
7.6K
被引数:
7.2W

机构

B
Beijing University of Chemical Technology
学者数:
3.1W
论文数: 2.2W
被引数: 4.5W
N
national university of defense technology - china
学者数:
1.8W
论文数: 1.4W
被引数: 9
引用论文

引用论文

Experimental and Analytical Study on the Operation Characteristics of the AHAT System
err2012-03-01
err0
PREAI
errHidefumi Araki; Tomomi Koganezawa; Chihiro Myouren; Shinichi Higuchi; Toru Takahashi; Takashi Eta
err分享
err收藏
Natural Actor-Critic
err2008-03-01
err643
PREAI
errPeters, Jan; Schaal, Stefan
err分享
err收藏
Informing sequential clinical decision-making through reinforcement learning: an empirical study
err2010-12-22
err136
errOAAI
errShortreed, Susan M.; Laber, Eric; Lizotte, Daniel J.; Stroup, T. Scott; Pineau, Joelle; Murphy, Susan A.
err分享
err收藏
学者 查看更多内容