返回
Distributional generative adversarial imitation learning with reproducing kernel generalization
DOI:10.1016/j.neunet.2023.05.027.png)
摘要
En 中文
Generative adversarial imitation learning (GAIL) regards imitation learning (IL) as a distribution matching problem between the state-action distributions of the expert policy and the learned policy. In this paper, we focus on the generalization and computational properties of policy classes. We prove that the generalization can be guaranteed in GAIL when the class of policies is well controlled. With the capability of policy generalization, we introduce distributional reinforcement learning (RL) into GAIL and propose the greedy distributional soft gradient (GDSG) algorithm to solve GAIL. The main advantages of GDSG can be summarized as: (1) Q-value overestimation, a crucial factor leading to the instability of GAIL with off-policy training, can be alleviated by distributional RL. (2) By considering the maximum entropy objective, the policy can be improved in terms of performance and sample efficiency through sufficient exploration. Moreover, GDSG attains a sublinear convergence rate to a stationary solution. Comprehensive experimental verification in MuJoCo environments shows that GDSG can mimic expert demonstrations better than previous GAIL variants. & COPY; 2023 Elsevier Ltd. All rights reserved.
Keyword:
Generative adversarial imitation learning
Policy generalization
Computational properties
Distributional reinforcement learning
期刊
IF:
6.3
论文数:
7.8K
被引数:
3.0W
机构
引用论文
Study of bi-directional buck-boost converter topologies for application in electrical vehicle motor drives应用于电动汽车电机驱动的双向buck-boost变换器拓扑研究
Improve generated adversarial imitation learning with reward variance regularization
MACHINE LEARNING
IF2.9

