返回
A parallel scheduling algorithm for reinforcement learning in large state space
DOI:10.1007/s11704-012-1098-y.png)
摘要
En 中文
The main challenge in the area of reinforcement learning is scaling up to larger and more complex problems. Aiming at the scaling problem of reinforcement learning, a scalable reinforcement learning method, DCS-SRL, is proposed on the basis of divide-and-conquer strategy, and its convergence is proved. In this method, the learning problem in large state space or continuous state space is decomposed into multiple smaller subproblems. Given a specific learning algorithm, each subproblem can be solved independently with limited available resources. In the end, component solutions can be recombined to obtain the desired result. To address the question of prioritizing subproblems in the scheduler, a weighted priority scheduling algorithm is proposed. This scheduling algorithm ensures that computation is focused on regions of the problem space which are expected to be maximally productive. To expedite the learning process, a new parallel method, called DCS-SPRL, is derived from combining DCS-SRL with a parallel scheduling architecture. In the DCS-SPRL method, the subproblems will be distributed among processors that have the capacity to work in parallel. The experimental results show that learning based on DCS-SPRL has fast convergence speed and good scalability.
Keyword:
divide-and-conquer strategy
parallel schedule
scalability
large state space
continuous state space
期刊
IF:
4.6
论文数:
1.6K
被引数:
2.8K
机构
引用论文
Studies in the flexibility of macrocyclic ligands. Crystal and molecular structure of 2,13-dimethy1-3,6,9,12,18-pentaazabicyclo (12.3.1) octadeca-(18),14,16-triene-dichloroion (III) hexafluorophosphate
Polyhedron
IF0
Taking Minorities for Granted? Ethnic Density, Party Campaigning and Targeting Minority Voters in 2010 British General Elections将少数群体视为理所当然?——民族密度、政党竞选活动与针对少数族裔选民的目标策略:以2010年英国大选为例
Efficient exploration through active learning for value function approximation in reinforcement learning
NEURAL NETWORKS
IF6.3
没有更多内容

