Return
Generalization error for portable rewards in transfer imitation learning
DOI:10.1016/j.knosys.2024.112230.png)
Abstract
En 中文
The reward transfer paradigm in transfer imitation learning (TIL) leverages the reward learned via inverse reinforcement learning (IRL) in the source environment to re-optimize a policy in the target environment. This paradigm aims to enhance sample efficiency while existing literature experimentally shows a wide range of transfer effects from satisfactory to poor. Thus, the theoretical measurement of its transfer effect still needs further studies. In this paper, we present the measurable generalization error bound between the true underlying policy and the learned policy via the reward transfer paradigm. Specifically, the error bound is separable and labeled into three categories: the inevitable training error (only depending on the assigned function classes and the number of samples), the optimizable training error (relying on the extracted reward) and the environmental deviation error. Further, we show that the environmental deviation error dominates the transfer effect for the large environmental difference, while the optimizable training error has significant contributions for the opposite case. From this theoretical standpoint, we predict transfer effects and propose reasonable reward transfer plans for different scenarios.
Keywords:
Transfer imitation learning
Reward transfer
Measurable generalization error bound
Transfer effects evaluation
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W

