arrow
返回

Quantum Imitation Learning

delete2024-10-01
delete0
delete
OA
AI
Z
Zhihao Cheng
K
Kaining Zhang
沈力 封面图
沈力 (Li Shen)
D
Dacheng Tao *
DOI:10.1109/TNNLS.2023.3275075delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Despite remarkable successes in solving various complex decision-making tasks, training an imitation learning (IL) algorithm with deep neural networks (DNNs) suffers from the high computation burden. In this work, we propose quantum imitation learning (QIL) with a hope to utilize quantum advantage to speed up IL. Concretely, we develop two QIL algorithms, quantum behavioural cloning (Q-BC) and quantum generative adversarial imitation learning (Q-GAIL). Q-BC is trained with a negative log-likelihood loss in an off-line manner that suits extensive expert data cases, whereas Q-GAIL works in an inverse reinforcement learning scheme, which is on-line and on-policy that is suitable for limited expert data cases. For both QIL algorithms, we adopt variational quantum circuits (VQCs) in place of DNNs for representing policies, which are modified with data re-uploading and scaling parameters to enhance the expressivity. We first encode classical data into quantum states as inputs, then perform VQCs, and finally measure quantum outputs to obtain control signals of agents. Experiment results demonstrate that both Q-BC and Q-GAIL can achieve comparable performance compared to classical counterparts, with the potential of quantum speed-up. To our knowledge, we are the first to propose the concept of QIL and conduct pilot studies, which paves the way for the quantum era.
Keyword:
Behavioral cloning (BC)
imitation learning (IL)
inverse reinforcement learning (IRL)
quantum IL (QIL)
variational quantum circuits (VQCs)

期刊

IEEE Transactions on Neural Networks and Learning Systems 封面图
IEEE Transactions on Neural Networks and Learning Systems
IF:
8.9
论文数:
7.6K
被引数:
7.2W

机构

U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
引用论文

引用论文

err分享
err收藏
Towards a wireless and fully-implantable ECoG system
err2013-06-01
err0
PREAI
errE. Tolstosheeva; J. Hoeffmann; J. Pistor; D. Rotermund; T. Schellenberg; D. Boll; T. Hertzberg; V. Gordillo-Gonzalez; S. Mandon; D. Peters-Drolshagen; M. Schneider; K. Pawelzik; A. Kreiter; S. Paul; W. Lang
err分享
err收藏
Gene-environment interactions in hypertension
err1999-01-01
err0
PREAI
errZdenka Pausova; Johanne Tremblay; Pavel Hamet
err分享
err收藏
DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
err2018-07-30
err322
errOAAI
errPeng, Xue Bin; Abbeel, Pieter; Levine, Sergey; van de Panne, Michiel
err分享
err收藏
The Role of Multiple Large Shareholders in the Choice of Debt Source
err2016-11-04
err0
PREAI
errSabri Boubaker; Wael Rouatbi; Walid Saffar
err分享
err收藏
学者 查看更多内容