1
Return

Exploration and Exploitation: A Study on Sample Efficiency in Reinforcement Learning With Multifaceted Curiosity Rewards and Adaptive Experience Replay Utilisation in Sparse Reward Environments

delete2026-07-30
delete0
delete
OA
AI
J
Jingyi Huang
G
GuiPeng Xi
M
Meixiu Lin
H
Haohui Zhang
B
Bo Li *
E
Evgeny Neretin
DOI:10.1049/cit2.70153delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
In recent years, the widespread application of deep reinforcement learning (DRL) in autonomous systems has highlighted the importance of achieving high sample efficiency under sparse reward conditions. To improve sample efficiency in sparse reward environments, this paper proposes a reinforcement learning framework built upon the Soft Actor Critic architecture, which integrates multifaceted curiosity rewards (MCR) and adaptive experience replay utilisation (AERU) (MCR-AERU SAC). MCR combines multi-level intrinsic motivational signals, such as state prediction error and model uncertainty, to provide rich exploration incentives, encouraging the controlled entity to deviate from existing trajectories and discover new high-reward behaviours. AERU dynamically adjusts the experience replay priorities based on posterior temporal difference error (TD error), focusing on utilising ‘partially successful’ transitions that are easily overlooked. The synergy between MCR and AERU enables the proposed framework to achieve an optimal balance between exploration and exploitation, significantly accelerating policy convergence and improving sample utilisation. Extensive experiments in complex dynamic training environments demonstrate that the proposed MCR-AERU SAC algorithm achieves up to 1.41 times the early-stage gain rate and a 55.56% improvement in task success rate compared to the HER-SAC baseline, demonstrating superior sample efficiency and excellent robustness in large-scale sparse reward environments.
Keywords:
adaptive experience replay
curiosity reward
exploration-exploitation trade-off
mixed exploration strategy
reinforcement learning sample efficiency
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

CAAI Transactions on Intelligence Technology cover
CAAI Transactions on Intelligence Technology
IF:
7.3
Papers:
649
Citations:
2.4K

Organization

N
northwestern polytechnical university
Scholars:
1.0W
Papers: 3.8K
Citations: 0
Moscow Aviation Institute cover
Moscow Aviation Institute
Scholars:
554
Papers: 340
Citations: 280
B
Beijing Institute of Spacecraft System Engineering
Scholars:
140
Papers: 87
Citations: 1
Cited Papers

Cited Papers

Citing Papers

Citing Papers