arrow
Return

Sample-efficient backtrack temporal difference deep reinforcement learning

delete2025-10-09
delete0
PRE
AI
刘琪 (Qi Liu)
P
Pengbin Chen
K
Ke Lin
K
Kaidong Zhao
J
Jinliang Ding
李彦杰 cover
李彦杰 (Yanjie Li)
DOI:10.1016/j.knosys.2025.114613delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Deep reinforcement learning algorithms often require large amounts of training data, particularly in robotic control tasks. To address this limitation, we propose a sample-efficient backtrack temporal difference learning method that enhances target state-action ( Q) value estimation. The proposed method dynamically prioritizes transitions based on their proximity to terminal states using backtrack sampling weights. This prioritization mechanism yields more accurate target Q-values, thereby improving the overall Q-value estimation precision. Furthermore, our analysis uncovers a novel link between curriculum learning and Bellman equation optimization. The proposed method is versatile, applicable to both discrete and continuous action spaces, and readily integrable with off-policy actor-critic algorithms. Extensive experiments show that the proposed method considerably reduces Q-value approximation errors and outperforms baselines across diverse benchmarks, achieving a 28 % performance improvement in four discrete action-space tasks and a 78 % gain in four continuous control tasks.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

N
Northeastern University
Scholars:
2.4W
Papers: 1.5W
Citations: 3.0W