arrow
Return

Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm

delete2014-07-02
delete34
delete
OA
AI
R
Róbert Busa‐Fekete *
B
Balázs Szörényi
P
Paul Weng
程玮玮 (Weiwei Cheng)
E
Eyke Hüllermeier
DOI:10.1007/s10994-014-5458-8delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
We introduce a novel approach to preference-based reinforcement learning, namely a preference-based variant of a direct policy search method based on evolutionary optimization. The core of our approach is a preference-based racing algorithm that selects the best among a given set of candidate policies with high probability. To this end, the algorithm operates on a suitable ordinal preference structure and only uses pairwise comparisons between sample rollouts of the policies. Embedding the racing algorithm in a rank-based evolutionary search procedure, we show that approximations of the so-called Smith set of optimal policies can be produced with certain theoretical guarantees. Apart from a formal performance and complexity analysis, we present first experimental studies showing that our approach performs well in practice.
Keywords:
Preference learning
Reinforcement learning
Evolutionary direct policy search
Racing algorithms
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.7K
Citations:
3.4W

Organization

P
Philipps University Marburg
Scholars:
1.3W
Papers: 1.0W
Citations: 10
S
Sorbonne Universite
Scholars:
6.2W
Papers: 4.5W
Citations: 605