arrow
返回

Improve generated adversarial imitation learning with reward variance regularization

delete2022-01-03
delete8
delete
OA
AI
Y
Yi-Feng Zhang *
F
Fan-Ming Luo
Y
Yang Yu
DOI:10.1007/s10994-021-06083-7delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Imitation learning aims at recovering expert policies from limited demonstration data. Generative Adversarial Imitation Learning (GAIL) employs the generative adversarial learning framework for imitation learning and has shown great potentials. GAIL and its variants, however, are found highly sensitive to hyperparameters and hard to converge well in practice. One key issue is that the supervised learning discriminator has a much faster learning speed than the reinforcement learning generator, making the generator gradient vanishing. Although GAIL is formulated as a zero-sum adversarial game, the ultimate goal of GAIL is to learn the generator, thus the discriminator should play the role more like a teacher rather than a real opponent. Therefore, the learning of the discriminator should consider how the generator could learn. In this paper, we disclose that enhancing the gradient of the generator training is equivalent to increase the variance of the fake reward provided by the discriminator output. We thus propose an improved version of GAIL, GAIL-VR, in which the discriminator also learns to avoid generator gradient vanishing through regularization of the fake rewards variance. Experiments in various tasks, including locomotion tasks and Atari games, indicate that GAIL-VR can improve the training stability and imitation scores.
Keyword:
Imitation learning
Reinforcement learning
Generative adversarial model
Discriminator reward

期刊

Machine Learning 封面图
Machine Learning
IF:
2.9
论文数:
2.7K
被引数:
3.4W

机构

N
nanjing university
学者数:
7.8W
论文数: 5.6W
被引数: 87
引用论文

引用论文

Musculoskeletal extremity pain in Danish school children – how often and for how long? The CHAMPS study-DK
err2017-11-25
err0
errOAAI
errSigne Fuglkjær; Jan Hartvigsen; Niels Wedderkopp; Eleanor Boyle; Eva Jespersen; Tina Junge; Lisbeth Runge Larsen; Lise Hestbæk
err分享
err收藏
Effects of riluzole on N-methyl-d-aspartate-induced tyrosine phosphorylation in the rat hippocampus
err2001-06-01
err0
PREAI
errAgnès Peyclit; Hawa Keita; Philippe Juvin; Pascal Derkinderen; Fanny Jardinaud; Danielle Rouellé; Jorge Boczkowski; Jean-Marie Desmonts; Jean-Antoine Girault; Jean Mantz
err分享
err收藏
Three-Dimensional Printed Modeling of an Arteriovenous Malformation Including Blood Flow
err2016-06-01
err0
PREAI
errJayesh P. Thawani; Jared M. Pisapia; Nickpreet Singh; Dmitriy Petrov; James M. Schuster; Robert W. Hurst; Eric L. Zager; Bryan A. Pukenas
err分享
err收藏
Is Host Metabolism the Missing Link to Improving Cancer Outcomes?
err2020-08-19
err0
errOAAI
errChristopher M. Wright; Anuradha A. Shastri; Emily Bongiorno; Ajay Palagani; Ulrich Rodeck; Nicole L. Simone
err分享
err收藏
The role of lockups in takeover contests
err2008-01-03
err0
errOAAI
errYeon‐Koo Che; Tracy R. Lewis
err分享
err收藏
Cks1 is degraded via the ubiquitin‐proteasome pathway in a cell cycle‐dependent manner
err2003-10-30
err0
PREAI
errTakayuki Hattori; Kyoko Kitagawa; Chiharu Uchida; Toshiaki Oda; Masatoshi Kitagawa
err分享
err收藏
没有更多内容