arrow
Return

Double actor-critic with TD error-driven regularization in reinforcement learning

delete2025-11-19
delete0
delete
OA
AI
H
Haohui Chen
Z
Zhiyong Chen
A
Aoxiang Liu
W
Wentuo Fang
DOI:10.1016/j.neunet.2025.108323delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
To obtain better value estimation in reinforcement learning, we propose a novel algorithm based on the double actor-critic framework with temporal difference error-driven regularization, abbreviated as TDDR. TDDR employs double actors, with each actor paired with a critic, thereby fully leveraging the advantages of double critics. Additionally, TDDR introduces an innovative critic regularization architecture. Compared to classical deterministic policy gradient-based algorithms that lack a double actor-critic structure, TDDR provides superior estimation. Moreover, unlike existing algorithms with double actor-critic frameworks, TDDR does not introduce any additional hyperparameters, significantly simplifying the design and implementation process. The convergence of the proposed TDDR to the optimal value is analyzed under random updating and simultaneous updating patterns. Extensive experiments on various tasks, including MuJoCo and Box2D, demonstrate that TDDR performs competitively against 13 algorithms, including both benchmarks and state-of-the-art methods. It also achieves statistically significant performance gains across several environments.
Keywords:
Reinforcement learning
Actor-critic
Double actors
Critic regularization
Temporal difference
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

C
Central South University
Scholars:
10.0W
Papers: 7.2W
Citations: 10.9W
T
the university of newcastle
Scholars:
563
Papers: 267
Citations: 3