返回
A temporal difference method for multi-objective reinforcement learning
DOI:10.1016/j.neucom.2016.10.100.png)
摘要
En 中文
This work describes MPQ-learning, an algorithm that approximates the set of all deterministic non dominated policies in multi-objective Markov decision problems, where rewards are vectors and each component stands for an objective to maximize. MPQ-learning generalizes directly the ideas of Q-learning to the multi-objective case. It can be applied to non-convex Pareto frontiers and finds both supported and unsupported solutions. We present the results of the application of MPQ-learning to some benchmark problems. The algorithm solves successfully these problems, so showing the feasibility of this approach. We also compare MPQ-learning to a standard linearization procedure that computes only supported solutions and show that in some cases MPQ-learning can be as effective as the scalarization method. (C) 2017 Elsevier B.V. All rights reserved.
Keyword:
Reinforcement learning
Multi-objective optimization
MOMDPs
Q-leaming
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Empirical evaluation methods for multiobjective reinforcement learning algorithms
MACHINE LEARNING
IF2.9
Hypervolume indicator and dominance reward based multi-objective Monte-Carlo Tree Search
MACHINE LEARNING
IF2.9
Reinforcement learning based sensing policy optimization for energy efficient cognitive radio networks
NEUROCOMPUTING
IF6.5
没有更多内容

