arrow
返回

A learning search algorithm with propagational reinforcement learning

delete2021-03-22
delete0
PRE
AI
W
Wei Zhang *
DOI:10.1007/s10489-020-02117-0delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
When reinforcement learning with a deep neural network is applied to heuristic search, the search becomes a learning search. In a learning search system, there are two key components: (1) a deep neural network with sufficient expression ability as a heuristic function approximator that estimates the distance from any state to a goal; (2) a strategy to guide the interaction of an agent with its environment to obtain more efficient simulated experience to update the Q-value or V-value function of reinforcement learning. To date, neither component has been sufficiently discussed. This study theoretically discusses the size of a deep neural network for approximating a product function of p piecewise multivariate linear functions. The existence of such a deep neural network with O(n + p) layers and O(dn + dnp + dp) neurons has been proven, where d is the number of variables of the multivariate function being approximated, is the approximation error, and n = O(p + log(2)(pd/)). For the second component, this study proposes a general propagational reinforcement-learning-based learning search method that improves the estimate h(.) according to the newly observed distance information about the goals, propagates the improvement bidirectionally in the search tree, and consequently obtains a sequence of more accurate V-values for a sequence of states. Experiments on the maze problems show that our method increases the convergence rate of reinforcement learning by a factor of 2.06 and reduces the number of learning episodes to 1/4 that of other nonpropagating methods.
Keyword:
Machine learning
Heuristic search
Reinforcement learning
Deep neural network
Deep learning
Learning search

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

N
northeastern university - china
学者数:
3.1W
论文数: 2.7W
被引数: 37
引用论文

引用论文

err分享
err收藏
Reinforced cuckoo search algorithm-based multimodal optimization
err2019-01-02
err48
PREAI
errThirugnanasambandam, Kalaipriyan; Prakash, Sourabh; Subramanian, Venkatesan; Pothula, Sujatha; Thirumal, Vengattaraman
err分享
err收藏
err
IF0
err
err0
errOAAI
err
err分享
err收藏
Effect of two prophylaxis methods on marginal gap of Cl Vresin-modified glass-ionomer restorations
err2016-03-01
err0
errOAAI
errSoodabeh Kimyai; Fatemeh Pournaghi-Azar; Mehdi Daneshpooy; Mehdi Abed Kahnamoii; Farnaz Davoodi
err分享
err收藏
学者 查看更多内容