arrow
Return

Reward-optimizing learning using stochastic release plasticity

delete2025-08-14
delete0
PRE
AI
Y
Yuhao Sun
W
Wantong Liao†
J
Jinhao Li
X
Xinche Zhang†
G
Guan Wang
Z
Zhiyuan Ma
宋森 cover
宋森 (Sen Song) *
DOI:10.3389/fncir.2025.1618506delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Synaptic plasticity underlies adaptive learning in neural systems; offering a biologically plausible framework for reward-driven learning. However; a question remains: how can plasticity rules achieve robustness and effectiveness comparable to error backpropagation? In this study; we introduce Reward-Optimized Stochastic Release Plasticity (RSRP); a learning framework where synaptic release is modeled as a parameterized distribution. Utilizing natural gradient estimation; we derive a synaptic plasticity learning rule that effectively adapts to maximize reward signals. Our approach achieves competitive performance and demonstrates stability in reinforcement learning; comparable to Proximal Policy Optimization (PPO); while attaining accuracy comparable with error backpropagation in digit classification. Additionally; we identify reward regularization as a key stabilizing mechanism and validate our method in biologically plausible networks. Our findings suggest that RSRP offers a robust and effective plasticity learning rule; especially in a discontinuous reinforcement learning paradigm; with potential implications for both artificial intelligence and experimental neuroscience.
Keywords:
synaptic plasticity
reinforcement learning
reward optimization
natural gradient
biologically plausible networks

Journal

Frontiers in Neural Circuits cover
Frontiers in Neural Circuits
IF:
3
Papers:
1.7K
Citations:
4.6K

Organization

T
tsinghua university
Scholars:
11.7W
Papers: 10.0W
Citations: 137