返回
Reinforcement learning method based on sample regularization and adaptive learning rate for AGV path planning
DOI:10.1016/j.neucom.2024.128820.png)
摘要
En 中文
This paper proposes the proximal policy optimization (PPO) method based on sample regularization (SR) and adaptive learning rate (ALR) to address the issues of limited exploration ability and slow convergence speed in Autonomous Guided Vehicle (AGV) path planning using reinforcement learning algorithms in dynamic environments. Firstly, the regularization term based on empirical samples is designed to solve the bias and imbalance issues of training samples, and the sample regularization is added to the objective function to improve the policy selectivity of the PPO algorithm, thereby increasing the AGV's exploration ability during the training process in the working environment. Secondly, the Fisher information matrix of the Kullback-Leibler (KL) divergence approximation and the KL divergence constraint term are exploited to design the policy update mechanism based on the dynamically adjustable adaptive learning rate throughout training. The method considers the geometric structure of the parameter space and the change of the policy gradient, aiming to optimize parameter update direction and enhance convergence speed and stability of the algorithm. Finally, the AGV path planning scheme based on reinforcement learning is established for simulation verification and comparations in two-dimensional raster map and Gazebo 3D simulation environment. Simulation results verify the feasibility and superiority of the proposed method applied to the AGV path planning problem.
Keyword:
Reinforcement learning
AGV
Path planning
Sample regularization
Adaptive learning rate
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
暂无机构信息
引用论文
Anti-conflict AGV path planning in automated container terminals based on multi-agent reinforcement learning基于多agent强化学习的自动化集装箱码头防冲突AGV路径规划
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Adaptive formation control of leader-follower mobile robots using reinforcement learning and the Fourier series expansion
ISA TRANSACTIONS
IF6.5
Fabrication of Microgel-Reinforced Hydrogels via Vat Photopolymerization通过vat光聚合制备微凝胶增强水凝胶
ACS MACRO LETTERS
IF5.2

