arrow
Return

Dynamic preference inference network: Improving sample efficiency for multi-objective reinforcement learning by preference estimation

delete2024-11-01
delete0
PRE
AI
Y
Yang Liu
Y
Ying Zhou
Z
Ziming He
J
Jingchen Li *
DOI:10.1016/j.knosys.2024.112512delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Multi-objective reinforcement learning (MORL) addresses the challenge of optimizing policies in environments with multiple conflicting objectives. Traditional approaches often rely on scalar utility functions, which require predefined preference weights, limiting their adaptability and efficiency. To overcome this, we propose the Dynamic Preference Inference Network (DPIN), a novel method designed to enhance sample efficiency by dynamically estimating the trajectory decision preference of the agent. DPIN leverages a neural network to predict the most favorable preference distribution for each trajectory, enabling more effective policy updates and improving overall performance in complex MORL tasks. Extensive experiments in various benchmark environments demonstrate that DPIN significantly outperforms existing state-of-the-art methods, achieving higher scalarized returns and hypervolume. Our findings highlight DPIN's ability to adapt to varying preferences, reduce sample complexity, and provide robust solutions in multi-objective settings.
Keywords:
Multi-objective reinforcement learning
Sample efficiency
Reinforcement learning

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

N
Northwestern Polytechnical University
Scholars:
4.6W
Papers: 3.7W
Citations: 5.3W
B
Z
zhejiang university
Scholars:
17.5W
Papers: 12.0W
Citations: 152
researcher View more organizations