arrow
Return

LinFa-Q: Accurate Q-learning with linear function approximation

delete2025-01-01
delete0
PRE
AI
Z
Zhechao Wang
Q
Qiming Fu *
J
Jianping Chen
Q
Quan Liu
Y
You Lu
H
Hongjie Wu
F
Fuyuan Hu
DOI:10.1016/j.neucom.2024.128654delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Although Q-learning has achieved remarkable success in some practical cases, it often suffers from the overestimation problem in stochastic environments, which is commonly viewed as a shortcoming of Q-learning. Overestimated values are introduced by estimations on the next state Q value, which is well-known as the maximization bias. In this paper, we propose a more accurate method for estimating the Q value by the Q value decomposition and re-evaluation with similar samples based on linear function approximation. Specifically, we reform the parameterized incremental update formula of Q-learning and also demonstrate that the new formula is equivalent to the original one. Moreover, we propose a new parameterized incremental update formula of Q-learning to address the overestimation problem and present the more accurate computing method, which can be used in problems with continuous state spaces and stochastic environments. Experimentally, when compared with Doubly Bounded Q-learning and other Q-learning based methods, the new algorithm has more than 31% improvement of performance in Mountain Car and Cart Pole. Furthermore, the algorithm is robust to the learning rate and its memory capacity. Finally, the practical applicability of our algorithm is discussed through an analysis of time consumption.
Keywords:
Reinforcement learning
Q-learning
Maximization bias
Q value decomposition
Linear function approximation

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

S
suzhou university of science & technology
Scholars:
4.9K
Papers: 4.8K
Citations: 4
S
soochow university - china
Scholars:
5.1W
Papers: 3.5W
Citations: 82