arrow
Return

Discrete-time mean-variance strategy based on reinforcement learning

delete2026-03-13
delete0
PRE
AI
S
Si Zhao
X
Xun Li *
X
Xiangyu Cui *
DOI:10.1080/01605682.2026.2627279delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This article studies a discrete-time mean-variance model based on reinforcement learning. Compared with its continuous-time counterpart in the literature, the discrete-time model makes more general assumptions about the asset’s return distribution. Using entropy to measure the cost of exploration, we derive the optimal investment strategy, whose density function is Gaussian type. Additionally, we design the corresponding reinforcement learning algorithm. Both simulation experiments and empirical analysis indicate that our discrete-time model exhibits better applicability when analysing real-world data than the continuous-time model.
Keywords:
Mean-variance
reinforcement learning
model exploration

Journal

Journal of the Operational Research Society cover
Journal of the Operational Research Society
IF:
2.7
Papers:
389
Citations:
9.2K

Organization

S
shanghai university of finance and economics
Scholars:
219
Papers: 167
Citations: 4
E
east china normal university
Scholars:
3.0W
Papers: 2.1W
Citations: 25
T
the hong kong polytechnic university
Scholars:
4.1K
Papers: 2.3K
Citations: 0
S
Shanghai University of Finance and Economics
Scholars:
2.0K
Papers: 2.5K
Citations: 4.0K
researcher View more organizations