返回
Optimal Actor-Critic Policy With Optimized Training Datasets
DOI:10.1109/TETCI.2022.3140375.png)
摘要
En 中文
Actor-critic (AC) algorithms are known for their efficacy and high performance in solving reinforcement learning problems, but they also suffer from low sampling efficiency. An AC based policy optimization process is iterative and needs to access the agent-environment to evaluate and update the policy by rolling out the policy, collecting rewards and states (i.e. samples), and learning from them. It ultimately requires a huge number of samples to learn an optimal policy. To improve sampling efficiency, we propose a strategy to optimize the training dataset that contains significantly less samples collected from the AC process. The dataset optimization is made of a best episode only operation, a policy parameter-fitness model, and a genetic algorithm module. The optimal policy network trained by the optimized training dataset exhibits superior performance compared to many contemporary AC algorithms in controlling autonomous dynamical systems. Evaluation on standard benchmarks shows that the method improves sampling efficiency, ensures faster convergence to optima, and is more data-efficient than its counterparts.
Keyword:
Training
Optimization
Approximation algorithms
Genetic algorithms
Convergence
Standards
Machine learning algorithms
Actor critic
reinforcement learning
policy optimization
genetic algorithm
training dataset optimization
期刊
I
IF:
6.5
论文数:
1.4K
被引数:
4.5K
机构
引用论文
Off-Policy Actor-Critic Structure for Optimal Control of Unknown Systems With Disturbances具有扰动的未知系统最优控制的非策略参与者-批评者结构
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Fabrication of porous hollow γ-Al2O3 nanofibers by facile electrospinning and its application for water remediation静电纺丝法制备多孔中空 γ-Al2O3纳米纤维及其在水体修复中的应用

