arrow
Return

Generalization error for portable rewards in transfer imitation learning

delete2024-09-01
delete0
PRE
AI
Y
Yirui Zhou
王磊 cover
王磊 (Lei Wang)
M
Mengxiao Lu
Z
Zhiyuan Xu
唐建 (Jian Tang)
Y
Yangchun Zhang
彭亚新 cover
彭亚新 (Yaxin Peng) *
DOI:10.1016/j.knosys.2024.112230delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The reward transfer paradigm in transfer imitation learning (TIL) leverages the reward learned via inverse reinforcement learning (IRL) in the source environment to re-optimize a policy in the target environment. This paradigm aims to enhance sample efficiency while existing literature experimentally shows a wide range of transfer effects from satisfactory to poor. Thus, the theoretical measurement of its transfer effect still needs further studies. In this paper, we present the measurable generalization error bound between the true underlying policy and the learned policy via the reward transfer paradigm. Specifically, the error bound is separable and labeled into three categories: the inevitable training error (only depending on the assigned function classes and the number of samples), the optimizable training error (relying on the extracted reward) and the environmental deviation error. Further, we show that the environmental deviation error dominates the transfer effect for the large environmental difference, while the optimizable training error has significant contributions for the opposite case. From this theoretical standpoint, we predict transfer effects and propose reasonable reward transfer plans for different scenarios.
Keywords:
Transfer imitation learning
Reward transfer
Measurable generalization error bound
Transfer effects evaluation

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

M
Midea
Scholars:
228
Papers: 155
Citations: 1
S
shanghai university
Scholars:
3.9W
Papers: 2.7W
Citations: 52