arrow
返回

Q-Managed: A new algorithm for a multiobjective reinforcement learning

delete2021-04-01
delete7
PRE
AI
A
Adrião Duarte Dória Neto
DOI:10.1016/j.eswa.2020.114228delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Multi-objective reinforcement learning (MORL) involves the use of reinforcement learning techniques to address problems with multiple objectives, conflicting or not. Among the main techniques used to treat this class of problems, we saw that they are limited by some factors, such as the Pareto Front shape and computational cost. This paper proposes a new iterative algorithm based on the single-policy approach, called Q-Managed. We use a hybrid multi-objective optimization (MOO) method that provides the mathematical guarantee that all policies belonging to the Pareto Front can be found, regardless of whether it is concave, convex or a mixture of both. Another important aspect that is worth mentioning is that its simplicity and performance are from a single-policy algorithms. To validate our proposal, we use the traditional MORL benchmarks and with different configurations of the Pareto Front. The Q-Managed shows success in finding all the optimal policies in all environments, surpassing all the single-policy algorithms in the literature in terms of policy quality. Based on the used benchmarks, its effectiveness can also be equated to the best multi-policy algorithms. The hypervolume metric was used to compare the quality of the policies found by our algorithm with those found in the state of the art. Extensions for non-episodic environments and stochastic transition functions are also introduced.
Keyword:
Multiobjective reinforcement learning
epsilon-Constraint
Q-Learning
Pareto dominance
Single-policy approach
Hypervolume
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
2.9W
被引数:
10.2W

机构

Universidade Federal do Rio Grande do Norte 封面图
Universidade Federal do Rio Grande do Norte
学者数:
9.7K
论文数: 5.4K
被引数: 5.2K
引用论文

引用论文

err分享
err收藏
Deficient cerebellar long-term depression and impaired motor learning in mGluR1 mutant mice
errCell
IF0
err1994-10-01
err0
PREAI
errAtsu Alba; Masanobu Kano; Chong Chen; Mark E. Stanton; Gregory D. Fox; Karl Herrup; Theresa A. Zwingman; Susumu Tonegawa
err分享
err收藏
Softmax exploration strategies for multiobjective reinforcement learning
err2017-11-01
err37
errOAAI
errVamplew, Peter; Dazeley, Richard; Foale, Cameron
err分享
err收藏
Empirical evaluation methods for multiobjective reinforcement learning algorithms
err2010-12-22
err162
errOAAI
errVamplew, Peter; Dazeley, Richard; Berry, Adam; Issabekov, Rustam; Dekker, Evan
err分享
err收藏
学者 查看更多内容